Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “software resilience”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Technical Resilience Navigator

The Technical Resilience Navigator (TRN) is a systematic approach to identifying vulnerabilities with energy and water systems, and prioritizing solutions that reduce risk. The ultimate outcome of the TRN is a set of actionable resilience solutions that address the site's most important gaps in resilience and enhance the ability to maintain mission continuity. The TRN is designed to step users through this planning process, providing a framework to: assign roles and responsibilities; collect and document information and data; document key inputs and outputs for each of the TRN modules (Site Level Planning, Baseline Development, Risk Assessment, Solution Development, and Solution Prioritization); document prioritized list of resilience solutions; and track progress through the entire process. Currently "software as a service" at the listed website.

Rotondo, Julia↗

Near Term Reliability and Resilience: Revisiting Resilience Metrics for the Electric Grid

This report presents the metrics employed in the Near-Term Reliability and Resilience (NTRR) project to study the inter-dependencies between electric and natural gas infrastructures, particularly under challenging conditions. These metrics were developed and applied to evaluate the reliability and resilience of the electric grid and natural gas systems in near-term scenarios (within the next 10 years) involving extreme weather events and major supply disruptions. The report defines the metrics, explains how they are calculated, and describes the process by which they are used to evaluate reliability and resilience across simulated scenarios. It also demonstrates how the resilience metrics integrate with other project activities and summarizes the software tools deployed to calculate and visualize the results.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Resilient Space Habitat Design Using Safety Controls

Space habitats will involve a complex and tightly coupled combination of hardware, software, and humans, while operating in challenging environments that pose many risks, both known and unknown. It will not be possible to design habitats that are immune to failure, nor will it be possible to foresee all possible failures. Rather than aiming for designs where ―failure is not an option,‖ habitats must be resilient to disruptions. We propose an approach to resilient design for space habitats based on the concept of safety controls from system safety engineering. We model disruptions using a state-and-trigger approach, where the space habitat is in one of three distinct states at each time instance: nominal, hazardous, or accident. We use safety controls as ways of preventing a system from entering or remaining in a hazardous or accident state. We develop a safety control option space for the habitat, from which designers can select the set of safety controls that best meet resilience, performance, and other system goals. The safety control option space is likely to be large, accordingly, we design a database that links safety controls to the applicable states and triggers. We demonstrate our approach on the early design stage of a Martian space habitat.

Safety↗

Accelerating computing for the future electric grid (CRADA Final Report)

As a participant in the Cyclotron Road Lab-Embedded Entrepreneurship Program (LEEP), Vellex Computing, Inc. has successfully validated the "Vellex Computing Stack," a breakthrough Analog Neural Computer (ANC) specifically designed for high-performance edge optimization. This project achieved critical milestones in mixed-signal circuit stability and software-hardware co-design, directly addressing national priorities in semiconductor resiliency. The success of this work is deeply rooted in the support from the Cyclotron Road LEEP, which provided the essential "hard tech" runway—funding, mentorship, and access to Lawrence Berkeley National Laboratory’s world-class characterization facilities—allowing Vellex to overcome the "Valley of Death" often faced by deep-tech hardware startups. By leveraging LBNL’s advanced testing infrastructure, Vellex was able to rigorously benchmark the ANC architecture against state-of-the-art digital solutions, a feat that would have been resource-prohibitive independently. This collaboration has not only advanced American leadership in analog computing but has also matured Vellex’s technology to a stage ripe for private sector commercialization.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Resilient Information Architecture Platform for Smart Grid (RIAPS)

A number of emerging trends will substantially alter the operation and control of the electric grid over the next several decades. These trends include ensuring resiliency under severe weather events, increasing integration of renewable electricity generation, supporting changing electricity demand patterns, and the improving cost effectiveness of distributed energy resources. To address these challenges, the future “Smart Grid” management will need to transition from centralized to coordinated distributed control paradigm. Reliable operation of the Smart Grid depends on distributed intelligence realized through software applications that run on distributed computing devices attached to the power system to collect data and collaboratively manage resources. However, much of the existing software for Smart Grid-enabled devices is either proprietary or developed with custom solutions, which limits interoperability among the heterogeneous devices and hinders the ability to manage system-level reliability, security, and resiliency requirements. Additionally, this approach makes Smart Grid applications hard to maintain, evolve, verify, and replace; resulting in high development and deployment costs. Further development of the Smart Grid requires a reusable software base-layer to move from hard-coded functionality to a plug-and-play architecture capable of managing system-level objectives and constraints in addition to providing consistent common services across heterogeneous devices and applications. Vanderbilt University, in collaboration with North Carolina State University and Washington State University has developed a foundation ‘software platform’ for developing and deploying robust, reliable, effective and secure software applications for the Smart Grid. The Resilient Information Architecture Platform for the Smart Grid (RIAPS) provides core services for building effective and powerful smart grid applications. It offers unique services for real-time data dissemination, fault tolerance, and coordination across apps distributed over the network.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Towards Software Bill of Materials in the Nuclear Industry

Large, modern industrial facilities often incorporate thousands of digital assets in their operational technology. Regulated facilities, such as nuclear power plants (NPPs), maintain robust cybersecurity and configuration management programs that often use bills of materials (BOMs) for these assets, including make, model, and version of hardware, firmware, and software. However, these BOMs typically capture only first- or second-tier information provided by the original equipment manufacturer (OEM). Unfortunately, as indicated by the increasing number and sophistication of software supply chain attacks, this level of detail is insufficient for identifying all the potential vulnerabilities and risks in software applications. Software BOMs (SBOMs) provide detailed enumeration of components and dependencies within the product or devices, including firmware. SBOMs can be combined with vulnerability data sources and vendor vulnerability attestations to improve vulnerability management and enable rapid identification of affected components when new software vulnerabilities are discovered. Ideally, SBOMs are created by the OEM prior to installation. However, since this practice is not yet commonplace and since NPPs are typically slow to adopt new technology, most NPPs do not incorporate SBOMs into their asset or configuration management programs. Fortunately, SBOMs can be generated by NPPs on existing digital assets to provide further insight into risk management decisions. This report provides an overview of the current SBOM ecosystem and recommends guidance on how to get started in a “crawl, walk, run” manner to develop and implement a sustainable SBOM program for digital assets in an NPP.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Holistic Measurement Driven Resilience: Combining Operational Fault and Failure Measurements and Fault Injection for Quantifying Fault Detection, Propagation and Impact. Final report

For HPC systems to date, application resilience to faults and failures has been accomplished by the brute- force method of checkpoint/restart, which allows an application to make forward progress in the face of system and application faults, errors, and failures independent of root cause or end result. It has remained the primary resilience mechanism because we lack a way to identify faults and anticipate consequences early enough to take meaningful mitigating action. However, checkpoint/restart implementations put a tremendous burden on system resources and on the applications themselves and is becoming less feasible at scale. Because we have not yet operated at scales at which checkpoint/restart fails to provide forward progress, despite increasing costs, vendors have had little motivation to provide the instrumentation necessary for early identification of faults and failures. However, as we move from petascale to exascale, component mean time to failure (MTTF) will render the existing techniques ineffectual and/or too expensive. Furthermore, fault recovery mechanisms such as failover and/or error correction introduce performance inconsistency. Instrumentation allowing early indication of problems and tools to enable use of such information by systems, operating systems, and applications offer an alternative, more scalable and less costly solution. In the HMDR project, we built on our experience and expertise developed and accumulated over years of research on design, monitoring, measurement, and assessment of resilient computing systems. Analysis of field data on the current and past generations of extreme-scale systems revealed several challenges that, if not addressed in increasingly larger and more complex systems, may hinder the effectiveness of future exascale computing systems. Specifically, i) file systems and interconnects in current-generation large-scale systems already operate at the margins of resiliency, including consistent performance, and may not scale to larger deployments; ii) automated, software-based failover mechanisms are frequently inadequate and can introduce wider failures, such that failures during recovery may lead to system/application failures, including system-wide outages; and iii) silent data corruption represents a critical fault mode and will require efficient detection mechanisms if next-generation applications are to take full advantage of exascale hardware. To address the above challenges, we assembled a team of world-renowned experts in resilient extreme- scale computing from the University of Illinois (Electrical and Computer Engineering, Computer Science, and NCSA), SNL, LANL, NERSC, and Cray. Our team includes representatives from centers that house many of the largest HPC resources in the world, both today and over the coming years. The team has a unique track record of research in i) system and application failure characterization based on the analysis of field data, ii) data-driven design of fault/error detection mechanisms, and iii) experimental characterization of system/application resiliency. The team includes system owners/operators who provide continuous data collection and access and ensure installation of appropriate analysis tools.

97 MATHEMATICS AND COMPUTING↗

Formally Verified ZTA Requirements for OT/ICS Environments with Isabelle/HOL

The clean energy transformation includes the integration of distributed energy resources with the power grid, which has led to a substantial increase in the complexity of power grids infrastructure and the underlying operational technology environment. Power grids infrastructure represents an operational technology environment that has become a system of systems, integrating heterogeneous devices which are both software-and hardware-intensive; as a result, there are increasing demands to exploit advances in the commodity of software-hardware infrastructures to improve energy systems requirements such as cybersecurity and resilience. In such a setting, system requirements at different levels mix, which leads to vulnerabilities and undesirable outcomes. The use of formal methods to characterize and prove system requirements removes ambiguity, increases automation, and provides high levels of assurance and reliability. In this paper, we contribute a methodology and a framework for the system-level verification of zero trust architecture requirements in operational technology environments. We define a formal specification for the core functionalities of operational technology environments, the corresponding invariants, and security proofs. Of particular note is our modular approach for the formal verification of asynchronous interactions in operational technology environments. The formal specification and the proofs have been mechanized using the interactive theorem proving environment Isabelle/HOL.

formal methods↗

MFC 5.0: An exascale many-physics flow solver

Many problems of interest in engineering, medicine, and the fundamental sciences rely on high-fidelity flow simulation, making performant computational fluid dynamics solvers a mainstay of the open-source software community. Previous work MFC 3.0 was made a published, documented, and open-source solver via Bryngelson et al. Comp. Phys. Comm. (2021) with numerous physical features, numerical methods, and scalable infrastructure. MFC 5.0 is a significant update to MFC 3.0, featuring a broad set of well-established and novel physical models and numerical methods, as well as the introduction of GPU and APU (or superchip) acceleration. Here, we exhibit state-of-the-art performance and ideal scaling on the first two exascale supercomputers, OLCF Frontier and LLNL El Capitan. Combined with MFC’s single-accelerator performance, MFC achieves exascale computation in practice, and achieved the largest-to-date public CFD simulation at 200 trillion grid points as a 2025 ACM Gordon Bell Prize finalist. New physical features include the immersed boundary method, N-fluid phase change, Euler–Euler and Euler–Lagrange sub-grid bubble models, fluid-structure interaction, hypo- and hyper-elastic materials, chemically reacting flow, two-material surface tension, magnetohydrodynamics (MHD), and more. Numerical techniques now represent the current state-of-the-art, including general relaxation characteristic boundary conditions, WENO variants, Strang splitting for stiff sub-grid flow features, and low Mach number treatments. Weak scaling to tens of thousands of GPUs on OLCF Summit and Frontier and LLNL El Capitan achieves efficiencies within 5% of ideal to over 90% of their respective system sizes. Strong scaling results for a 16-times increase in device count show parallel efficiencies over 90% on OLCF Frontier. MFC’s software stack has undergone further improvements, including continuous integration, which ensures code resilience and correctness through over 300 regression tests; metaprogramming, which reduces code length while maintaining performance portability; and code generation for computing chemical reactions

Computational fluid dynamics↗

Autonomous Inverter Controls for Resilient and Secure Grid Operation: Vector Control Design for Grid Forming

The project addresses both fundamental and practical challenges of GFM/GFL inverter control for the power grids with high inverter based resources (IBRs) penetration. A data- driven modeling technique is applied to accurately model dynamics of PWM inverters, including electromagnetic-transient (EMT). Systematic and integrative designs of grid- forming (GFM) and grid-following (GFL) primary controls are developed to guarantee system performance under either normal or abnormal operating conditions without violating constraints. This modeling and control framework provides black-start capability in case of an outage without relying on rotating generators, and its secondary control is also shown to enhance resilience against cyber-physical attacks.

14 SOLAR ENERGY↗

Overview of NASA's Air Traffic Management - eXploration (ATM-X) Project

Projected increases in new vehicle types, new missions, and the continual growth in traditional (e.g., airlines, general aviation) aviation will require changes to the current air traffic system, particularly to accommodate the desire of operators to be more involved in air traffic decisions. To address these challenges, the National Airspace System needs to undergo a transformation to a more scalable, flexible, user-focused system that addresses safety and security requirements and resiliency for current and new users. A system designed to integrate modular software services, provided by users, third parties and government for air traffic management functions, will be scalable and more easily allow modernization and for collaboration between users and service providers. ATM-X is responding to NASA's pivot towards integrating projected new, diverse entrants into the NAS, while also leveraging NASA's prior ATM achievements that continue to improve traditional airspace operations. This project is a two-phased approach to conduct research and focused evaluations to assess the feasibility of a service-based approach and to identify critical design considerations to enable airspace access for new entrants, integrated with current traditional operations. Phase 1 research will be conducted to determine what is needed to reach the ATM-X goals based on specific use-cases to enable large-scale, passenger-carrying Urban Air Mobility operations in a metroplex environment, and also to improve traditional operations in the Northeast Region leveraging mature NASA technologies. Some of these evaluations will be conducted in simulations and field activities. Phase 2 will build upon Phase 1 towards more defined, focused research and field demonstrations in real-world environments to integrate multiple elements of a scalable, service-based ATM-X concept.

air traffic management↗

Idaho National Laboratory Energy Cybersecurity Programs Update

This brief presentation provides a status update on three Idaho National Laboratory energy cybersecurity programs of particular interest to NERC Reliability and Security Technical Committee annual in-person Security Groups summit. Public information on the following three programs is included: Cybersecurity for Operational Technology Environments (CyOTE™) program Cyber-Informed Engineering (CIE) Cyber Testing for Resilient Industrial Control Systems (CyTRICS) program, and associated high-level information on the Energy Software Bill of Materials POC, Executive Order 14017, and the Energy Cyber Sense Act

24 POWER TRANSMISSION AND DISTRIBUTION↗

JUSTIFI: Software for Improving Performance Objectives via Energy Efficiency

With growing energy supply concerns and rising costs, energy efficiency is a critical component of industrial energy resilience and competitiveness by directly reducing energy operating costs. Energy efficiency projects in manufacturing also yield valuable benefits to other key metrics, such as improved quality, reduced maintenance costs, improved safety, decreased pollution, and enhanced productivity. However, it is difficult to receive approval for energy efficiency projects, so implementation rates are low, even when meeting capital project payback period criteria. The inclusion and quantification of non-energy benefits (NEBs) in the decision-making process for energy efficiency projects can improve the overall financial payback period while demonstrating a positive impact on the firm's key performance metrics and business strategy. Despite their significant financial and strategic value, NEBs are rarely factored into decision-making due to lack of tools to effectively identify and quantify them. Therefore, a comprehensive and integrative approach is needed for the rapidly evolving energy landscape. To address these challenges, through funding from U.S. Department of Energy, our new assessment methodology integrates common continuous improvement six sigma concepts, such as the DMAIC process, and a protocol of guiding questions, into energy efficiency assessments to identify NEBs. We have also developed open-source software, JUSTIFI, to guide users through this process, data collection, and quantification. It is designed to be used concurrently with DOE energy system analysis software suite, MEASUR. Our methodology and tools inform energy assessors, firm engineering, decision makers, and workforce seeking to increase energy resilience and to maximize benefits aligned with performance metrics.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

DSS-SimPy-RL (Open-DSS and SimPy based Cyber-Physical RL environment) [SWR-23-29]

Recently, numerous data-driven approaches to control an electric grid using machine learning techniques have been investigated. With the advancement of reinforcement learning (RL) based techniques, gradually the conventional optimization based solvers are being replaced with RL approach where there is uncertainty in the environment such as renewable generation or cyber system emulation. However, to train an agent efficiently, it requires numerous interactions with an environment to learn the best policies. There are numerous RL environments for the power systems based on some well-known simulators, similarly there are environment for communication domains. While majority of the cyber emulators are based in an UNIX environment, the power simulators are based in the Windows-based operating system, the generation of cyber-physical mixed domain RL environment has been challenging. Existing co-simulation methods are efficient but resource and time intensive to generate large scale data set for training RL agents. Hence, this software focuses on development and validation of a mixed domain RL environment using Open DSS for the physical side and leverages a discrete event simulator python package, SimPy, for cyber-side emulation which is Operating Systems agnostic. Further utilizing this software co-simulation and training RL agents for re-routing based resilient control for network reconfiguration and volt-var control in power distribution feeder are performed.

Sahu, Abhijeet↗

sup3ruhi (Super Resolution for Renewable Resource Data and Urban Heat Islands) [SWR-25-05]

Urban heat is a growing concern, particularly in dense metropolitan areas where high temperatures increase the risk of heat-related illness and drive energy expenses for cooling. Estimating the effects of urban heat remains a challenge due to limitations in describing the built environment, computational constraints, and the need for high-resolution data. This software presents open-source, computationally efficient machine learning methods that enhance the accuracy of urban temperature estimates compared to historical reanalysis data. Models trained using this software have been applied to urban microclimates in Los Angeles and Seattle showing greater accuracy and less bias when compared to low-resolution reanalysis datasets like ERA5 and even when compared to high-resolution mesoscale numerical weather models like WRF with an urban canopy model. Initial findings highlight how machine learning can support urban heat resilience planning by enabling improved assessments of local heat islands, mitigation strategies, and their energy implications. This software is an extension of (sup3r). This software supports the following publication: Buster, Grant, et al. Tackling Extreme Urban Heat: A Machine Learning Approach to Assess the Impacts of Climate Change and the Efficacy of Climate Adaptation Strategies in Urban Microclimates. arXiv:2411.05952, arXiv, 8 Nov. 2024. arXiv.org, https://doi.org/10.48550/arXiv.2411.05952. And has related public data records available at: Buster, Grant, Cox, Jordan, Benton, Brandon, and King, Ryan. Super-Resolution for Renewable Resource Data and Urban Heat Islands (Sup3rUHI). United States: N.p., 16 Oct, 2024. Web. https://data.openei.org/submissions/6220.

Buster, Grant [National Renewable Energy Laborator↗

JUSTIFI: Software for Improving Performance Objectives via Energy Efficiency

With growing energy supply concerns and rising costs, energy efficiency is a critical component of industrial energy resilience and competitiveness by directly reducing energy operating costs. Energy efficiency projects in manufacturing also yield valuable benefits to other key metrics, such as improved quality, reduced maintenance costs, improved safety, decreased pollution, and enhanced productivity. However, it is difficult to receive approval for energy efficiency projects, so implementation rates are low, even when meeting capital project payback period criteria. The inclusion and quantification of non-energy benefits (NEBs) in the decision-making process for energy efficiency projects can improve the overall financial payback period while demonstrating a positive impact on the firm's key performance metrics and business strategy. Despite their significant financial and strategic value, NEBs are rarely factored into decision-making due to lack of tools to effectively identify and quantify them. Therefore, a comprehensive and integrative approach is needed for the rapidly evolving energy landscape. To address these challenges, through funding from U.S. Department of Energy, our new assessment methodology integrates common continuous improvement six sigma concepts, such as the DMAIC process, and a protocol of guiding questions, into energy efficiency assessments to identify NEBs. We have also developed open-source software, JUSTIFI, to guide users through this process, data collection, and quantification. It is designed to be used concurrently with DOE energy system analysis software suite, MEASUR. Our methodology and tools inform energy assessors, firm engineering, decision makers, and workforce seeking to increase energy resilience and to maximize benefits aligned with performance metrics.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Resilient, Rural, and Revolutionary: Salisbury Square's Direct-Current Affordable Microgrid Community: Preprint

The technology to interconnect buildings with a dedicated direct-current (DC) power distribution network is in place today; what is missing is a turnkey approach to designing a DC microgrid - and the business models allowing such systems to be deployed, owned, and operated at scale. To close this gap, the Salisbury Square Development Team, comprising clean-energy experts, has engineered a resilient community DC microgrid for an affordable housing community in Randolph, Vermont. Ten single-family, occupant-owned residences and 12 multifamily rental units will share locally generated and stored solar energy via a DC power distribution bus capable of operating during extended grid outages. With a DC power distribution network in place, each home will be equipped with high-efficiency DC lighting and appliances, operating alongside alternating current (AC) appliances, even during an islanded mode of operation. To obtain a comprehensive understanding of what is possible and achievable, the Team collaborated with the local utility, regulatory agencies, a national laboratory, energy-as-service providers, and vendors. The collaborators evaluated microgrid typologies, business models, and energy modeling, and analyzed electrification and resilience. Further, the Team applied the URBANoptTM (Urban Renewable Building and Neighborhood optimization, NREL 2022) software development kit (SDK) to Salisbury Square's single-family and multifamily buildings to validate workflows and identify needs for advanced capability. This paper addresses the barriers to entry, scalability, and impact on residents and system ownership. It also examines the analysis that informed the design and engineering of the DC microgrid and the opportunities to streamline the process.

DER↗

Contingency Software in Autonomous Systems: Technical Level Briefing

Contingency management is essential to the robust operation of complex systems such as spacecraft and Unpiloted Aerial Vehicles (UAVs). Automatic contingency handling allows a faster response to unsafe scenarios with reduced human intervention on low-cost and extended missions. Results, applied to the Autonomous Rotorcraft Project and Mars Science Lab, pave the way to more resilient autonomous systems.

autonomous systems↗