Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “software resilience”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

OpenSAMPL: An Open Source Library for Timing and Synchronization Measurements and Analytics

Today's power grid operators are implementing timing and synchronization solutions that provide resilience to Global Navigation Satellite System (GNSS) vulnerabilities. These vendor-specific solutions often come with additional software applications that are designed to monitor that vendor's synchronization performance data. However, resilient timing architectures often resulting in multi-vendor solutions, including approaches that blend terrestrial clocks with space-based subscription services. In such an environment, collecting, analyzing, and visualizing data from a variety of sources within a single platform was heretofore not possible. To address this need, the US Department of Energy's Center for Alternative Synchronization and Timing (CAST) developed OpenSAMPL, the Open Synchronized Analytics and Monitoring Platform, an open-source Python framework for processing, loading, and observing clock measurement data from distributed devices. OpenSAMPL enables the ingestion of diverse clock-probe sources into a scalable time-series database and applies robust analytics. OpenSAMPL currently supports two vendor data pipelines, and will be extended to more in the near future, enabling seamless monitoring of a variety of timing and synchronization devices in a common environment.

Grant, Josh [ORNL] (ORCID:0000000163475060)↗

Energy-Efficient and Resilient Infrastructure: Simulation, Validation, and Installation

Advanced, high-performance computing at the National Renewable Energy Laboratory (NREL) has enabled access to vast data resources with cutting-edge software techniques to understand, design, plan for, and maintain energy-efficient and resilient infrastructure. We have focused on cities and airports, but the technology we have developed will easily translate to seaports, inland ports, military installations, or other complex and large-scale energy-intensive systems. We can digitally simulate and explore current and future scenarios to make datadriven decisions for optimizing advanced energy systems, transportation and building operations, infrastructure planning and expansion, and battery storage to guide short- and long-term investments, electrification strategies, and integration of new technologies.

Athena↗

Rapid Evaluation and Response to Impacts on Critical End-Use Loads Following Natural Hazard-Driven Power Outages: A Modular and Responsive Geospatial Technology

The disparate nature of data for electric power utilities complicates the emergency recovery and response process. The reduced efficiency of response to natural hazards and disasters can extend the time that electrical service is not available for critical end-use loads, and in extreme events, leave the public without power for extended periods. This article presents a methodology for the development of a semantic data model for power systems and the integration of electrical grid topology, population, and electric distribution line reliability indices into a unified, cloud-based, serverless framework that supports power system operations in response to extreme events. An iterative and pragmatic approach to working with large and disparate datasets of different formats and types resulted in improved application runtime and efficiency, which is important to consider in real time decision-making processes during hurricanes and similar catastrophic events. This technology was developed initially for Puerto Rico, following extreme hurricane and earthquake events in 2017 and 2020, but is applicable to utilities around the world. Given the highly abstract and modular design approach, this technology is equally applicable to any geographic region and similar natural hazard events. In addition to a review of the requirements, development, and deployment of this framework, technical aspects related to application performance and response time are highlighted.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Faster-than-real-time Simulation with Demonstration for Resilient DER Integration

The US electric grid is facing operational, stability, and security challenges. Transmission system operators need some measure of visibility into distribution system renewable generation. Distribution system generation needs to support transmission system voltage. The grid is experiencing an expansion in measurement systems. How to take full advantage of this expansion and defend against attacks, both cyber and physical, poses additional challenges. The Faster-than-real-time Simulation with demonstration for Resilient DER Integration project set out to do the following: a. Flatten the voltage profile through the feeders and system for cost saving and voltage stabilization needs. b. Increase the amount of intermittent distributed energy resources (IDERs) that could be deployed on a utility feeder and provide 100% or more energy needed for the demands on that feeder, and c. based on an accurate model (Digital Twin) of the utilities system, be able to detect any abnormalities on the utilities distribution system. To manage the voltage and increase IDER penetration (a,b), Graph Trace Analysis is employed in a time-series, optimal power flow to coordinate the time-varying feedback control setpoints of a distribution feeder’s utility control devices. Under the coordinated control are a Load Tap Changing Transformer, a voltage regulator, and five switched capacitor banks. The feeder serves over 2000 customers, the feeder secondaries are modeled, and the feeder has 2.3 MW of PV generation, corresponding to a 17.4% penetration of PV generation. The feeder model has over 12,000 components, where every customer load bus and PV generator are modeled. The accuracy of the power flow solution is compared against historical meter voltage measurements, the improvement in conservation voltage reduction energy savings as a function of the coordinated control desired voltage profile is investigated, and the increase in PV penetration of the coordinated control over the existing control is presented. To achieve improved control performance while observing system operation constraints, bellwether Advanced Metering Infrastructure (AMI) voltage measurements are used to adjust the desired voltage profile used by the optimal power flow analysis. To detect and alleviate or negate attacks or failures on the distribution and transmission utility grids (c) the grid needs to be resilient and self-healing. In this project software was designed to do just that. At the center of the software is an Integrated System Model (ISM) that spans from transmission to secondary distribution. The ISM is employed in real-time abnormality detection, voltage stability forecasting, and multi-mode control. Testing results are presented for: 1—attacks on utility infrastructure; 2—energy savings from optimal control; 3—distribution system control response during a low voltage transmission system event; 4—cyber-attacks on PV inverters, where physical inverters are used in hard-ware-in-the-simulation-loop studies. Contributions of this work include real-time analysis that spans from three-phase transmission through secondary distribution; an approach for detecting abnormalities that employs measurements from three independent measurement systems; and a multi-mode distribution system control that responds to cyber-attacks, physical attacks, equipment failures, and transmission system needs.

Integrated System Model, Graph Trace Analysis, Adv↗

Mission aware cyber safe mode for spacecraft

Mission-Aware Cyber Safe Mode (MACSM) is spacecraft resilience architecture that enables autonomous containment of software-level cyber intrusions while maintaining control authority and mission continuity. It is analogous to traditional safe modes that preserve vehicle survival by shutting down non-essential subsystems in response to faults or environmental stress.

97 MATHEMATICS AND COMPUTING↗

Framework for Assessing Impact of Wave-Powered Desalination on Resilience of Coastal Communities

Coastal communities face unique challenges in maintaining continuous service from critical infrastructure. This research advances capabilities for evaluating the impact of using wave energy to desalinate water on the resilience of coastal communities. The study focuses on the feasibility of using wave energy conversion to provide drinking water to communities in need and applying resilience metrics to quantify its impact on the community. To assess the feasibility of wave-powered desalination, this research couples the open-source software Wave Energy Converter SIMulator (WEC-Sim) and Water Network Tool for Resilience (WNTR). This research explores variations in both the wave resource (location, seasonality, and duration) and the ability to maintain drinking water service during a disruption scenario by applying the simulation framework to three case studies, which are based on communities in Puerto Rico. The simulation framework provides a contextualized assessment of the ability of wave-powered desalination to improve the resilience of coastal communities, which can serve as a methodology for future studies seeking the integration of wave-powered desalination with water distribution systems.

16 TIDAL AND WAVE POWER↗

CEEP (Cyber-Energy Emulation Platform) [SWR-20-102]

NREL's Cyber-Energy Emulation Platform (CEEP) provides the capability to realize cyber-energy security and resilience through automation and orchestration of virtualized systems and software defined networks for the electric grid. CEEP enables testing and validation of grid-security and -control methodologies as the grid evolves to include smart technologies/systems, such as virtualization and containerization of grid components, software defined networking, simulation and co-simulation frameworks, and hardware in the loop. CEEP is a modular system that can be distributed and deployed across different hardware infrastructure sizes and network architectures. For example, CEEP can visualize, emulate, and/or coordinate the Smart-Grid Network Visualization, Intrusion Detection, and Network Healing system. Using CEEP, intrusion-detection and network-self-healing solutions can be deployed at grid control centers, within secure private clouds, and in cyber-energy appliances.

Vaughan, Evan↗

Cyber Energy Emulation Platform (CEEP) [SWR-20-102]

NREL's Cyber-Energy Emulation Platform (CEEP) provides the capability to realize cyber-energy security and resilience through automation and orchestration of virtualized systems and software defined networks for the electric grid. CEEP enables testing and validation of grid-security and -control methodologies as the grid evolves to include smart technologies/systems, such as virtualization and containerization of grid components, software defined networking, simulation and co-simulation frameworks, and hardware in the loop. CEEP is a modular system that can be distributed and deployed across different hardware infrastructure sizes and network architectures. For example, CEEP can visualize, emulate, and/or coordinate the Smart-Grid Network Visualization, Intrusion Detection, and Network Healing system. Using CEEP, intrusion-detection and network-self-healing solutions can be deployed at grid control centers, within secure private clouds, and in cyber-energy appliances.

Rivera, Joshua↗

Data and Tools for Energy Planning and Analysis

The National Renewable Energy Laboratory (NREL) creates widely used data and tools to facilitate energy system planning and analysis. These software tools have been developed for complex research problems and perfected over real-world applications and laboratory validations. Some tools are award winners, others are open-source data explorers, and all are rigorously designed to empower decision-makers with accurate and accessible information. This software selection shows how NREL resources can help stakeholders achieve a clean, just, and resilient energy transformation.

data↗

Technical Resilience Navigator

The Technical Resilience Navigator (TRN) is a systematic approach to identifying vulnerabilities with energy and water systems, and prioritizing solutions that reduce risk. The ultimate outcome of the TRN is a set of actionable resilience solutions that address the site's most important gaps in resilience and enhance the ability to maintain mission continuity. The TRN is designed to step users through this planning process, providing a framework to: assign roles and responsibilities; collect and document information and data; document key inputs and outputs for each of the TRN modules (Site Level Planning, Baseline Development, Risk Assessment, Solution Development, and Solution Prioritization); document prioritized list of resilience solutions; and track progress through the entire process. Currently "software as a service" at the listed website.

Rotondo, Julia↗

Near Term Reliability and Resilience: Revisiting Resilience Metrics for the Electric Grid

This report presents the metrics employed in the Near-Term Reliability and Resilience (NTRR) project to study the inter-dependencies between electric and natural gas infrastructures, particularly under challenging conditions. These metrics were developed and applied to evaluate the reliability and resilience of the electric grid and natural gas systems in near-term scenarios (within the next 10 years) involving extreme weather events and major supply disruptions. The report defines the metrics, explains how they are calculated, and describes the process by which they are used to evaluate reliability and resilience across simulated scenarios. It also demonstrates how the resilience metrics integrate with other project activities and summarizes the software tools deployed to calculate and visualize the results.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Accelerating computing for the future electric grid (CRADA Final Report)

As a participant in the Cyclotron Road Lab-Embedded Entrepreneurship Program (LEEP), Vellex Computing, Inc. has successfully validated the "Vellex Computing Stack," a breakthrough Analog Neural Computer (ANC) specifically designed for high-performance edge optimization. This project achieved critical milestones in mixed-signal circuit stability and software-hardware co-design, directly addressing national priorities in semiconductor resiliency. The success of this work is deeply rooted in the support from the Cyclotron Road LEEP, which provided the essential "hard tech" runway—funding, mentorship, and access to Lawrence Berkeley National Laboratory’s world-class characterization facilities—allowing Vellex to overcome the "Valley of Death" often faced by deep-tech hardware startups. By leveraging LBNL’s advanced testing infrastructure, Vellex was able to rigorously benchmark the ANC architecture against state-of-the-art digital solutions, a feat that would have been resource-prohibitive independently. This collaboration has not only advanced American leadership in analog computing but has also matured Vellex’s technology to a stage ripe for private sector commercialization.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Resilient Information Architecture Platform for Smart Grid (RIAPS)

A number of emerging trends will substantially alter the operation and control of the electric grid over the next several decades. These trends include ensuring resiliency under severe weather events, increasing integration of renewable electricity generation, supporting changing electricity demand patterns, and the improving cost effectiveness of distributed energy resources. To address these challenges, the future “Smart Grid” management will need to transition from centralized to coordinated distributed control paradigm. Reliable operation of the Smart Grid depends on distributed intelligence realized through software applications that run on distributed computing devices attached to the power system to collect data and collaboratively manage resources. However, much of the existing software for Smart Grid-enabled devices is either proprietary or developed with custom solutions, which limits interoperability among the heterogeneous devices and hinders the ability to manage system-level reliability, security, and resiliency requirements. Additionally, this approach makes Smart Grid applications hard to maintain, evolve, verify, and replace; resulting in high development and deployment costs. Further development of the Smart Grid requires a reusable software base-layer to move from hard-coded functionality to a plug-and-play architecture capable of managing system-level objectives and constraints in addition to providing consistent common services across heterogeneous devices and applications. Vanderbilt University, in collaboration with North Carolina State University and Washington State University has developed a foundation ‘software platform’ for developing and deploying robust, reliable, effective and secure software applications for the Smart Grid. The Resilient Information Architecture Platform for the Smart Grid (RIAPS) provides core services for building effective and powerful smart grid applications. It offers unique services for real-time data dissemination, fault tolerance, and coordination across apps distributed over the network.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Towards Software Bill of Materials in the Nuclear Industry

Large, modern industrial facilities often incorporate thousands of digital assets in their operational technology. Regulated facilities, such as nuclear power plants (NPPs), maintain robust cybersecurity and configuration management programs that often use bills of materials (BOMs) for these assets, including make, model, and version of hardware, firmware, and software. However, these BOMs typically capture only first- or second-tier information provided by the original equipment manufacturer (OEM). Unfortunately, as indicated by the increasing number and sophistication of software supply chain attacks, this level of detail is insufficient for identifying all the potential vulnerabilities and risks in software applications. Software BOMs (SBOMs) provide detailed enumeration of components and dependencies within the product or devices, including firmware. SBOMs can be combined with vulnerability data sources and vendor vulnerability attestations to improve vulnerability management and enable rapid identification of affected components when new software vulnerabilities are discovered. Ideally, SBOMs are created by the OEM prior to installation. However, since this practice is not yet commonplace and since NPPs are typically slow to adopt new technology, most NPPs do not incorporate SBOMs into their asset or configuration management programs. Fortunately, SBOMs can be generated by NPPs on existing digital assets to provide further insight into risk management decisions. This report provides an overview of the current SBOM ecosystem and recommends guidance on how to get started in a “crawl, walk, run” manner to develop and implement a sustainable SBOM program for digital assets in an NPP.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Holistic Measurement Driven Resilience: Combining Operational Fault and Failure Measurements and Fault Injection for Quantifying Fault Detection, Propagation and Impact. Final report

For HPC systems to date, application resilience to faults and failures has been accomplished by the brute- force method of checkpoint/restart, which allows an application to make forward progress in the face of system and application faults, errors, and failures independent of root cause or end result. It has remained the primary resilience mechanism because we lack a way to identify faults and anticipate consequences early enough to take meaningful mitigating action. However, checkpoint/restart implementations put a tremendous burden on system resources and on the applications themselves and is becoming less feasible at scale. Because we have not yet operated at scales at which checkpoint/restart fails to provide forward progress, despite increasing costs, vendors have had little motivation to provide the instrumentation necessary for early identification of faults and failures. However, as we move from petascale to exascale, component mean time to failure (MTTF) will render the existing techniques ineffectual and/or too expensive. Furthermore, fault recovery mechanisms such as failover and/or error correction introduce performance inconsistency. Instrumentation allowing early indication of problems and tools to enable use of such information by systems, operating systems, and applications offer an alternative, more scalable and less costly solution. In the HMDR project, we built on our experience and expertise developed and accumulated over years of research on design, monitoring, measurement, and assessment of resilient computing systems. Analysis of field data on the current and past generations of extreme-scale systems revealed several challenges that, if not addressed in increasingly larger and more complex systems, may hinder the effectiveness of future exascale computing systems. Specifically, i) file systems and interconnects in current-generation large-scale systems already operate at the margins of resiliency, including consistent performance, and may not scale to larger deployments; ii) automated, software-based failover mechanisms are frequently inadequate and can introduce wider failures, such that failures during recovery may lead to system/application failures, including system-wide outages; and iii) silent data corruption represents a critical fault mode and will require efficient detection mechanisms if next-generation applications are to take full advantage of exascale hardware. To address the above challenges, we assembled a team of world-renowned experts in resilient extreme- scale computing from the University of Illinois (Electrical and Computer Engineering, Computer Science, and NCSA), SNL, LANL, NERSC, and Cray. Our team includes representatives from centers that house many of the largest HPC resources in the world, both today and over the coming years. The team has a unique track record of research in i) system and application failure characterization based on the analysis of field data, ii) data-driven design of fault/error detection mechanisms, and iii) experimental characterization of system/application resiliency. The team includes system owners/operators who provide continuous data collection and access and ensure installation of appropriate analysis tools.

97 MATHEMATICS AND COMPUTING↗

Formally Verified ZTA Requirements for OT/ICS Environments with Isabelle/HOL

The clean energy transformation includes the integration of distributed energy resources with the power grid, which has led to a substantial increase in the complexity of power grids infrastructure and the underlying operational technology environment. Power grids infrastructure represents an operational technology environment that has become a system of systems, integrating heterogeneous devices which are both software-and hardware-intensive; as a result, there are increasing demands to exploit advances in the commodity of software-hardware infrastructures to improve energy systems requirements such as cybersecurity and resilience. In such a setting, system requirements at different levels mix, which leads to vulnerabilities and undesirable outcomes. The use of formal methods to characterize and prove system requirements removes ambiguity, increases automation, and provides high levels of assurance and reliability. In this paper, we contribute a methodology and a framework for the system-level verification of zero trust architecture requirements in operational technology environments. We define a formal specification for the core functionalities of operational technology environments, the corresponding invariants, and security proofs. Of particular note is our modular approach for the formal verification of asynchronous interactions in operational technology environments. The formal specification and the proofs have been mechanized using the interactive theorem proving environment Isabelle/HOL.

formal methods↗

MFC 5.0: An exascale many-physics flow solver

Many problems of interest in engineering, medicine, and the fundamental sciences rely on high-fidelity flow simulation, making performant computational fluid dynamics solvers a mainstay of the open-source software community. Previous work MFC 3.0 was made a published, documented, and open-source solver via Bryngelson et al. Comp. Phys. Comm. (2021) with numerous physical features, numerical methods, and scalable infrastructure. MFC 5.0 is a significant update to MFC 3.0, featuring a broad set of well-established and novel physical models and numerical methods, as well as the introduction of GPU and APU (or superchip) acceleration. Here, we exhibit state-of-the-art performance and ideal scaling on the first two exascale supercomputers, OLCF Frontier and LLNL El Capitan. Combined with MFC’s single-accelerator performance, MFC achieves exascale computation in practice, and achieved the largest-to-date public CFD simulation at 200 trillion grid points as a 2025 ACM Gordon Bell Prize finalist. New physical features include the immersed boundary method, N-fluid phase change, Euler–Euler and Euler–Lagrange sub-grid bubble models, fluid-structure interaction, hypo- and hyper-elastic materials, chemically reacting flow, two-material surface tension, magnetohydrodynamics (MHD), and more. Numerical techniques now represent the current state-of-the-art, including general relaxation characteristic boundary conditions, WENO variants, Strang splitting for stiff sub-grid flow features, and low Mach number treatments. Weak scaling to tens of thousands of GPUs on OLCF Summit and Frontier and LLNL El Capitan achieves efficiencies within 5% of ideal to over 90% of their respective system sizes. Strong scaling results for a 16-times increase in device count show parallel efficiencies over 90% on OLCF Frontier. MFC’s software stack has undergone further improvements, including continuous integration, which ensures code resilience and correctness through over 300 regression tests; metaprogramming, which reduces code length while maintaining performance portability; and code generation for computing chemical reactions

Computational fluid dynamics↗

Autonomous Inverter Controls for Resilient and Secure Grid Operation: Vector Control Design for Grid Forming

The project addresses both fundamental and practical challenges of GFM/GFL inverter control for the power grids with high inverter based resources (IBRs) penetration. A data- driven modeling technique is applied to accurately model dynamics of PWM inverters, including electromagnetic-transient (EMT). Systematic and integrative designs of grid- forming (GFM) and grid-following (GFL) primary controls are developed to guarantee system performance under either normal or abnormal operating conditions without violating constraints. This modeling and control framework provides black-start capability in case of an outage without relying on rotating generators, and its secondary control is also shown to enhance resilience against cyber-physical attacks.

14 SOLAR ENERGY↗