Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “research data management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Data for publication: "A fresh take: Seasonal changes in terrestrial freshwater inputs impact salt marsh hydrology and vegetation dynamics"

This data repository contains data associated with the manuscript "A fresh take: Seasonal changes in terrestrial freshwater inputs impact salt marsh hydrology and vegetation dynamics". This study was conducted at the Elkhorn Slough National Estuarine Research Reserve in Watsonville, California from October 2019 - June 2022. We sought to understand the role of shallow freshwater inputs from adjacent uplands on salt marsh hydrologic behavior and vegetation productivity. This dataset contains CSV files of the following: daily salt marsh subsurface water level and pore water conductivity, monthly vegetation survey measurements, soil core data. Estuary surface water level, conductivity, and local precipitation were downloaded from the National Estuarine Research Reserve System (Centralized Data Management Office) at https://cdmo.baruch.sc.edu/.

54 ENVIRONMENTAL SCIENCES↗

The Silencing of U.S. Campuses Following the COVID-19 Response: Evaluating Root Mean Square Seismic Amplitudes Using Power Spectral Density Data

In response to the COVID-19 global pandemic, many populated and active regions have become deserted and show significant reductions in their background seismicity, especially campuses across the United States (U.S.). Seismic sensors located in the vicinity of or within U.S. campuses show that anthropogenic seismic noise remains elevated during the ordinary, nonpandemic, academic year, only subduing during periods of recess (e.g., winter break). Here, we use power spectral density (PSD) data computed by the Incorporated Research Institutions for Seismology Data Management Center for quality assessment to calculate root mean square (rms) amplitude and analyze the effects of the COVID-19 school closures. We processed and analyzed PSD data for 46 seismic stations located within 50 m of a U.S. university or college. Results show that 42 campus stations show an overall rms drop following a statewide school closure.

58 GEOSCIENCES↗

New Architecture to Support Integration and Processing of Seismic Data from Heterogeneous Sources

The Geophysical Monitoring Program (GMP) at Lawrence Livermore National Lab (LLNL) maintains a database and supporting infrastructure for geophysical data used in support of the Nuclear Detonation Detection mission. This database includes data from multiple sources, many of which do not distribute data to the public or for which there is no automated means of access. For example, Figure 1 shows (left) the distribution of waveform data in our database by source. The Incorporated Research Institutions for Seismology Data Management Center (IRISDMC) is our major source of waveform data and those data may be retrieved at will using the Federated Digital Seismograph Networks FDSN web Application Programming Interface (API). However, the next 6 most important sources of waveform data have no or only limited automated access to waveforms. As Figure 1 (right) shows, it is very common for waveform records associate with an event in our database to come from two or more sources, and in some cases data come from 10 sources. This diversity of data sources drives our need for efficient and correct integration of metadata, parametric data, and waveform data.

58 GEOSCIENCES↗

Data Release Report for the Source Physics Experiment Phase II: Dry Alluvium Geology Experiments (DAG-1 through DAG-4), Nevada National Security Site

The Dry Alluvium Geology (DAG) project was Phase II of the Source Physics Experiment and consisted of a series of four chemical explosive tests conducted in the same source hole on the Nevada National Security Site. This hole is located at 37.1146°N and -116.0693°W, with a surface elevation of 1,285.2 meters (m) (4,216.5 feet [ft]) above sea level. The first test (DAG-1) was conducted on July 20, 2018, at 16:51:52.67838 Coordinated Universal Time (UTC). The explosive source for DAG-1 was nitromethane initiated by a small plastic-bonded explosive (PBX) charge, detonated at the depth of 385.0 m (1,263.2 ft) below ground surface. DAG-1 had a trinitrotoluene (TNT) equivalent yield of 0.908 metric tons (2,002 pounds [lbs]). DAG-2 was conducted on December 19, 2018, at 18:45:56.92115 UTC. This test was the largest in the series, with a TNT equivalent yield of 50.997 metric tons (112,429 lbs). The explosive source for DAG-2 was nitromethane initiated by a small PBX charge, detonated at the depth of 299.8 m (983.6 ft) below ground surface. DAG-3 was conducted on April 27, 2019, at 15:49:01.84183 UTC. The explosive source for this test was nitromethane initiated by a small PBX charge, detonated at the depth of 149.9 m (492.0 ft) below ground surface. DAG-3 had a TNT equivalent yield of 0.908 metric tons (2,002 lbs). The final DAG test (DAG-4) was conducted on June 22, 2019, at 21:06:19.87632 UTC. The explosive source for DAG-4 was nitromethane initiated by a small PBX charge, detonated at the depth of 51.6 m (169.3 ft) below ground surface. DAG-4 had a TNT equivalent yield of 10.357 metric tons (22,833 lbs). The four tests were recorded by an extensive set of instrumentation that included sensors both at near-field (less than 200 m) and far-field (200 m or greater) distances. The near-field instruments consisted of three-component (3C) accelerometers installed at various depths ranging from 51.6 to 385 m (169.3 to 1,263.1 ft) below ground surface in boreholes positioned around the source hole, and arrays of single-component and 3C accelerometers on the surface. The far-field network comprised a variety of seismic and acoustic sensors, including short-period geophones, broadband seismometers, and 3C accelerometers at distances of 200 m to 400 kilometers. In addition, the DAG-2, DAG-3, and DAG-4 explosions were recorded by a temporary array of 496 geophones arranged in a densely spaced grid pattern known as “Large N.” This report coincides with the release of these data for analysts and organizations that are not participants in this program. This report describes the four DAG tests and the various types of near-field, far-field, and other data that are available. Assembled data sets are accessible through: Incorporated Research Institutions for Seismology, Data Management Center 1408 NE 45th Street, Suite 201, Seattle, Washington 98105 USA. www.iris.washington.edu

58 GEOSCIENCES↗

An open source fast fluid dynamics model for data center thermal management

Although computational fluid dynamics (CFD) has been widely adopted to improve data center thermal management, the high computational demand limits its applications, such as multivariate optimal design and operation. Fast fluid dynamics (FFD), which has been applied for fast airflow simulation, shows great potential. However, few research applied FFD for optimal design and operation of data center thermal management. This research improves the FFD model for data centers and conducts a comprehensive evaluation and demonstration. First, the FFD model is improved by solving the advection and diffusion equations together using an upwind scheme instead of a semi-Lagrangian advection solver in the conventional FFD model. Second, new features for data centers are added, such as a pressure correction method to simulate plenum airflow and dynamic boundary conditions for IT racks. The new FFD model is first validated with two indoor environment cases and the results show that the new FFD model has slightly better overall prediction accuracy and faster speed compared to the conventional FFD model. It is also observed that both FFD models achieve acceptable accuracy, except for a few localized disparities with experimental data, which might be due to simplified handling of turbulence viscosity near the boundaries. Furthermore, validation with a real data center shows that the FFD model achieves a similar level of accuracy as CFD when compared to the experimental measurements with some level of uncertainties. It is then demonstrated for data center optimal design and operation, which saves 53.4–58.8% of annual energy while still meeting the thermal requirements. In conclusion, with a much faster speed and comparable accuracy compared to CFD, the FFD model parallelized on a graphics processing unit is promising for practical model-based data center early design and operation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Supply Chain Risk Management: Data Structuring

Supply chain risk management (SCRM) is an area of research that addresses both logistics concepts to maximize efficiency, reliability, and revenue as well as risk features, such as potential weak points, break points, and vulnerabilities within the supply chain. SCRM is used to find risks introduced at each node in a supply chain and how these risks can impact a company’s products, individuals, customers, and reputation. SCRM is a relatively new field, so standardized processes including data structuring are not fully documented. This paper explains the importance of a standard data structuring methodology and how it can enhance current SCRM efforts. Data ingest, structuring, and analysis are predominantly managed by humans. Automating some of the less complex steps can positively impact SCRM by allowing human analysts to focus on more strategic analyses. Types of data to be collected and structured are collected via publicly available information related to hardware, software, and corporate entities. After the data has been collected, the information is formatted in a specific manner, conforming to a schema, to allow for more effective and efficient ingest for further analysis. This paper outlines data structures used by Pacific Northwest National Laboratory for SCRM research and analysis purposes. These structures have been used for hundreds of analyses and have been successful in developing a common baseline. Data structuring is one of the first steps in data standardization, which will further mature and enhance the SCRM research area.

supply chain risk management, data structuring, re↗

Employing Technology to Enable Remote Research Charrettes as a Method for Engaging Industry and Uncovering Best Practices: A Novel Approach for a Post-COVID-19 World

Methods to collect data in construction engineering and management (CEM) research are evolving, informed by recent technological advancements. One such method is research charrettes that allow effective interactions and knowledge sharing between expert industry practitioners and academic researchers, all colocated in a single venue, enabling rich data collection and live communication. A pivot point in technological evolution occurred with the COVID-19 pandemic, forcing a global shift to remote work. Hence, planned in-person research charrettes had to shift to remote sessions, relying on virtual conferencing platforms and online data collection mechanisms. Technology-enabled charrettes have allowed the authors to collect significantly richer data sets and ensure a more diverse representation of participants, while saving tremendous amounts of time. With the continuing emergence of technological applications, the world might not go back to functioning fully in person. The authors believe remote research charrettes (RRCs) will still be used in a post-COVID-19 world because of their superior performance. This paper builds on a previous publication that described traditional research charrettes as a method to enhance CEM research a decade ago; it offers a significantly updated and improved RRC method based on the knowledge gained from transitioning a dozen in-person charrettes into RRCs. It also presents performance comparisons between RRCs and traditional charrettes by quantifying metrics indicating how RRCs are more time-efficient and cost-saving, harness more participants from more diverse locations, and enable the collection of richer data sets and four times more industry comments and expert feedback. This paper also provides guidance on the integration of technology with traditional research charrettes, hence contributing to the CEM body of knowledge.

42 ENGINEERING↗

Data as a Key Resource in Catalysis: A Community Account

The deployment of artificial intelligence (AI) is transforming the scientific fields central to interdisciplinary catalysis research. By enabling more effective use of data, AI (including simpler machine learning and data science tools) holds great promise for accelerating discoveries. However, progress has so far been modest, largely due to the lack of standardized, machine-readable, and openly shared catalysis data. This perspective, accounting for community insights emerging at conferences, analyses the underlying reasons for these challenges and proposes solutions to a future whereFAIR data management becomes an integral part of research in catalysis. In the short-term, we deem that mandatory FAIR data depositing prior to scientific publications along with consensualized top-down guidelines on data sharing powered by ease-to-use tools can make the necessary step change happen to catalyse data as key resource in our community.

36 - MATERIALS SCIENCE↗

Sample Identifiers and Metadata to Support Data Management and Reuse in Multidisciplinary Ecosystem Sciences

Physical samples are foundational entities for research across biological, Earth, and environmental sciences. Data generated from sample-based analyses are not only the basis of individual studies, but can also be integrated with other data to answer new and broader-scale questions. Ecosystem studies increasingly rely on multidisciplinary team-science to study climate and environmental changes. While there are widely adopted conventions within certain domains to describe sample data, these have gaps when applied in a multidisciplinary context. In this study, we reviewed existing practices for identifying, characterizing, and linking related environmental samples. We then tested practicalities of assigning persistent identifiers to samples, with standardized metadata, in a pilot field test involving eight United States Department of Energy projects. Participants collected a variety of sample types, with analyses conducted across multiple facilities. We address terminology gaps for multidisciplinary research and make recommendations for assigning identifiers and metadata that supports sample tracking, integration, and reuse. Furthermore, our goal is to provide a practical approach to sample management, geared towards ecosystem scientists who contribute and reuse sample data.

54 ENVIRONMENTAL SCIENCES↗

Quality Assurance Program Plan for SFR Metallic Fuel Data Qualification

This document contains an evaluation of the applicability of the current Quality Assurance Standards from the American Society of Mechanical Engineers Standard NQA-1 (NQA-1) criteria and identifies and describes the quality assurance process(es) by which attributes of historical, analytical, and other data associated with sodium-cooled fast reactor [SFR] metallic fuel will be evaluated. This process is being instituted to facilitate validation of data to the extent that such data may be used to support future licensing efforts associated with advanced reactor designs. The initial data to be evaluated under this program were generated during the US Integral Fast Reactor program between 1984-1994, where the data include, but are not limited to, research and development data and associated documents, test plans and associated protocols, operations and test data, technical reports, and information associated with past United States Nuclear Regulatory Commission reviews of SFR designs. It is recognized that managing the data generated by large research and development projects presents a significant challenge for retaining data integrity and availability. American Society of Mechanical Engineers Standard NQA-1 (NQA-1) 2008/2009a provides appropriate requirements for this plan.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Quality Assurance Program Plan for SFR Metallic Fuel Data Qualification

This document contains an evaluation of the applicability of the current Quality Assurance Standards from the American Society of Mechanical Engineers Standard NQA-1 (NQA-1) criteria and identifies and describes the quality assurance process(es) by which attributes of historical, analytical, and other data associated with sodium-cooled fast reactor [SFR] metallic fuel will be evaluated. This process is being instituted to facilitate validation of data to the extent that such data may be used to support future licensing efforts associated with advanced reactor designs. The initial data to be evaluated under this program were generated during the US Integral Fast Reactor program between 1984-1994, where the data include, but are not limited to, research and development data and associated documents, test plans and associated protocols, operations and test data, technical reports, and information associated with past United States Nuclear Regulatory Commission reviews of SFR designs. It is recognized that managing the data generated by large research and development projects presents a significant challenge for retaining data integrity and availability. American Society of Mechanical Engineers Standard NQA-1 (NQA-1) 2008/2009a provides appropriate requirements for this plan.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A Vision for Coupling Operation of US Fusion Facilities with HPC Systems and the Implications for Workflows and Data Management

The operation of large US Department of Energy (DOE) research facilities, like the DIII-D National Fusion Facility, results in the collection of complex multi-dimensional scientific datasets, both experimental and model-generated. In the future, it is envisioned that integrated data analysis coupled with large-scale high performance computing (HPC) simulations will be used to improve experimental planning and operation. Practically, massive data sets from these simulations provide the physics basis for generation of both reduced semi-analytic and machine-learning-based models. Storage of both HPC simulation datasets (generated from US DOE leadership computing facilities) and experimental datasets presents significant challenges. In this paper, we present a vision for a DOE-wide data management workflow that integrates US DOE fusion facilities with leadership computing facilities. Data persistence and long-term availability beyond the length of allocated projects is essential, particularly for verification and recalibration of artificial intelligence and machine learning (AI/ML) models. Because these data sets are often generated and shared among hundreds of users across multiple leadership computing facility centers, they would benefit from cross-platform accessibility, persistent identifiers (e.g. DOI, or digital object identifier), and provenance tracking. Here, the ability to handle different data access patterns suggests that a combination of low cost, high latency (e.g. for storing ML training sets) and high cost, low latency systems (e.g. for real-time, integrated machine control feedback) may be needed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

2019 Budget Request for the DOE Computational Science Graduate Fellowship (CSGF) Grant

The Department of Energy Computational Science Graduate Fellowship (DOE CSGF) is necessary to meet the continual challenging national workforce needs that arise as computational science and engineering problems continue to grow in scope and complexity. Computational science and engineering (CSE) is a multidisciplinary approach that uses scientific computing to solve practical problems methods and to supply technical tools across the scientific discovery spectrum. In particular, the DOE CSGF emphasizes high-performance computing (HPC) that enables CSE that advances science and engineering in directions important to the DOE and the economy in general. Over the past half-century, HPC has been an essential tool for DOE’s success. During this period, important missions, such as nuclear stockpile stewardship, have turned to HPC as an essential technology. Entire science disciplines, such as biology and cosmology, have been transformed through the augmentation of scientific observation via HPC. At government laboratories and in industry, DOE CSGF alumni are helping push traditional HPC boundaries while contributing to discoveries in high-energy physics, renewable energy, fusion-reactor design, additive manufacturing, nanomaterials for next-generation batteries and transistors, and turbine and advanced nuclear reactor modeling. In addition, HPC is used to address national health needs that will eventually point to cures both by helping cancer researchers manage and analyze huge troves of data, by simulating biological mechanisms, and by accelerating drug development — including continuing to rise to the challenge of pandemic-related research. A 2023 report from the ASCAC Subcommittee on American Competitiveness and Innovation to the ASCR office, “Can the United States Maintain Its Leadership in High-Performance Computing?” says of the Program, “The CSGF program provides a barometer for disciplines that will be of interest to future DOE computing.” An explosion in scientific and technological data has driven the need for increasingly sophisticated HPC to transform those data into scientific understanding. With access to more and more data and the proliferation of HPC, Machine Learning and Artificial Intelligence are experiencing a renaissance, complementing the now well-established use of computational simulation. Indeed, in its September 2020 subcommittee report on “AI/ML, Data Intensive Science and High-Performance Computing”, the DOE Advanced Scientific Computing Advisory Committee (ASCAC) explicitly called for a fellowship program to train computational and data scientists to tackle exascale and data-intensive computing challenges. This collaboration of empirical and theory-based modeling will increasingly inform federal policymakers whose decisions affect American society and future generations, and it requires highly skilled and intellectually agile computational scientists who can support the fast-moving DOE National Laboratory research environment. In fact, the DOE CSGF program has explicitly and consistently addressed this need.

97 MATHEMATICS AND COMPUTING↗

Integrated Energy-Water Data for Cross-Sector Resilience

This white paper focuses on the “energy-for-water” domain, addressing the urgent need for integrated, empirical data to support regional management, benchmarking, and research on improving efficiency and developing technologies for water and wastewater management systems. The costs and energy required for the supply, treatment, and distribution of water and wastewater lack a standard data collection mechanism and centralized database or storage infrastructure, limiting data-driven decision-making across interdependent infrastructure systems.

42 ENGINEERING↗

The Geothermal Data Repository: Ten Years of Supporting the Geothermal Industry with Open Access to Geothermal Data: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) is celebrating its tenth anniversary! Over the last decade it has grown from the simple idea of storing public data in a centralized location to a valuable tool at the center of the US geothermal scientific community and an integral part of the DOE Geothermal Technologies Office (DOE GTO) project management strategy. Researchers funded by the DOE GTO have contributed over 1,300 data submissions to the GDR. These data have been used to further advancements in geothermal science, economic analysis, exploration, research, development, and operational efficiency. The adoption of open data methodologies and a data management strategy that prioritizes universal open access and standardized, interoperable data have further increased the value of GDR data, making them available across a distributed network of data sharing partners and improving their utility to other industries and related fields, including material science and space exploration. Incorporating feedback from users has been critical to the GDRs success, allowing it to grow over the years to meet the evolving needs of the geothermal community. This paper will explore some of many changes that occurred throughout the GDRs tenure and the lessons learned along the way, as well as highlight some of the new features and recent improvements that been implemented to support innovation, reduce duplication of effort, and advance the geothermal industry as a whole.

accessibility↗

The Geothermal Data Repository: Ten Years of Supporting the Geothermal Industry with Open Access to Geothermal Data

The Department of Energy's (DOE) Geothermal Data Repository (GDR) is celebrating its tenth anniversary! Over the last decade it has grown from the simple idea of storing public data in a centralized location to a valuable tool at the center of the US geothermal scientific community and an integral part of the DOE Geothermal Technologies Office (DOE GTO) project management strategy. Researchers funded by the DOE GTO have contributed over 1,300 data submissions to the GDR. These data have been used to further advancements in geothermal science, economic analysis, exploration, research, development, and operational efficiency. The adoption of open data methodologies and a data management strategy that prioritizes universal open access and standardized, interoperable data have further increased the value of GDR data, making them available across a distributed network of data sharing partners and improving their utility to other industries and related fields, including material science and space exploration. Incorporating feedback from users has been critical to the GDR's success, allowing it to grow over the years to meet the evolving needs of the geothermal community. This paper will explore some of many changes that occurred throughout the GDRs tenure and the lessons learned along the way, as well as highlight some of the new features and recent improvements that been implemented to support innovation, reduce duplication of effort, and advance the geothermal industry as a whole.

access↗

Integrated Hydro-Terrestrial Modeling: Development of a National Capability

Water is one of our most important natural resources and is essential to our national economy and security. Multiple federal government agencies have mission elements that address national needs related to water. Each water-related agency champions a unique science and/or operational mission focused on advancing a portion of the nation’s ability to meet our water-related challenges. These diverse mission needs have engendered a rich and extensive base of water-related data and modeling capabilities. While useful for their intended purposes, these capabilities are not well integrated to address complex regional problems and overarching national problems. These major investments by a number of federal agencies, however, lay the foundation for an integrated hydro-terrestrial modeling and data infrastructure that will enhance knowledge, understanding, prediction, and management of the nation’s diverse water challenges. Creating a more seamless national hydro-terrestrial modeling and data capacity presents an enormous opportunity to advance operational as well as research capabilities leading to more effective water management. Advances are necessary not only in operational tools for forecasting but also in research to identify and resolve knowledge and data gaps that lead to unacceptable uncertainties in forecast outcomes. As such, close coordination across scientific, operational, and resource management communities is required. To this end, an interagency workshop on “Integrated Hydro-Terrestrial Modeling: Development of a National Capacity” was held at the National Science Foundation (NSF) headquarters in Alexandria, Virginia in September 2019, led jointly by the NSF, the U.S. Department of Energy (DOE) and the U.S. Geological Survey (USGS) and with broader interagency support provided through an interagency steering committee. This workshop provided a venue to bring together representatives of water-related agencies and their scientific partners (including university researchers) to initiate and refine a vision for a national Integrated Hydro-Terrestrial Modeling (IHTM) and data infrastructure and advance ideas towards its development. The workshop was designed to address three critical foci to advance the development of a national IHTM capacity: “Priority Water Challenges” around which to motivate and initiate development; Technical and methodological obstacles related to data and modeling; Organizational, structural, and cultural barriers that heretofore have impeded integration of capabilities across the federal and research landscapes. The following “Priority Water Challenge” domain areas represent targets for initiating development of the IHTM and were identified and selected in alignment with priorities of the administration’s Water Sub-Cabinet: (1) Nutrient loading, hypoxia, and harmful algal blooms; (2) Water availability in the western United States; and (3) Extreme weather-related water hazards. These water challenges span agency mission boundaries and encompass a broad range of geographies, complex system dynamics and feedbacks, and critical processes spanning hydrological, climatic, and biophysical systems as well as land-use/land-cover, agricultural, built infrastructure, societal, economic, and decisional environments. These three Priority Water Challenges cannot be fully addressed without leveraging complementary and synergistic capabilities across multiple agencies.

99 GENERAL AND MISCELLANEOUS↗

Development of Analysis Methods that Integrate Numeric and Textual Equipment Reliability Data

Within the Light Water Reactor Sustainability (LWRS) program, the Risk-Informed Systems Analysis (RISA) Pathway is performing collaborative research on the development and deployment of technologies designed to assist operating nuclear power plants (NPPs) to reduce operating costs improve plant reliability and availability. One of the RISA research areas is focusing on the development of methods and tools designed to optimize plant operations (e.g., maintenance/replacement schedules, optimal maintenance postures for plant structures, systems, and components [SSCs]) in a manner that is more cost effective than current approaches and makes better use of available SSC health data. The Risk-Informed Asset Management (RIAM) project targets this research area by creating a direct bridge between component equipment reliability (ER) data and system engineer decision making regarding maintenance activity scheduling and component aging management. In this respect, one challenge that NPP system engineers are facing is that the amount of ER data being continuously generated is not only extremely large in size, but it comes in different forms: textual (e.g., condition or maintenance reports) and numeric (e.g., generated by monitoring systems). All these data elements provide them with valuable insights and information regarding: 1) the discovery of anomalous behaviors or degradation trends, 2) the identification of the possible causes behind such behaviors/trends, and 3) the prediction of their direct consequences. However, several challenges have proved to be roadblocks to this process. While some of these challenges are technical in nature (i.e., data are often distributed over several physical servers/databases), others are conceptual in nature: data elements come in different formats (e.g., numeric or textual), and measured values have different scales (e.g., vibration spectra and oil temperature). The activities performed by the RIAM project during FY23 directly tackles the need to simultaneously integrate the analysis of ER data in all its forms, numeric and textual. Note that such task has never been performed before due to the complexity of the systems under consideration but, most importantly, because of the technical challenges behind the harmonization of ER data formats and the lack of adequate computational methods to analyze them. Our approach borrows ideas and concepts from the medical field where integration of several data sources is vital to assist medical practitioners to perform correct diagnosis and indicate optimal treatments. In our view a NPP asset is equivalent to a patient in a medical context. The main difference is the complexity of a human body is a magnitude more complex when compared to typical assets commonly present in NPPs (e.g., centrifugal pumps, or motor operated valves). This simplifies our first requirement when analyzing heterogenous ER data formats: to put data into “context”. Context is here intended as the additional piece of information that is needed by ER data analysis tools to understand what these data elements are referring to, i.e., which king of knowledge they are generating. In our context, this knowledge can be translated into models that capture the form and functional architecture of assets/systems, their dependencies, and how they interact. These models actually emulate the knowledge that that NPP system engineers possess about assets and systems; this is their key of success when analyzing ER data, their challenge is ability to handle large amount of data. Here, we employ model-based system engineering (MBSE) models of systems and assets to represent and capture their architecture and functional, i.e. cause-effect, relations. Then, ER data elements are processed by identifying first of all which elements of the developed MBSE elements they are referring to. For numeric ER data this task is fairly easy since it is possible to precisely pinpoint what MBSE elements the corresponding sensor are observing (e.g., bearing temperature of a centrifugal pump). Task is much harder for textual data since the information contained in issue or maintenance reports needs to “be understood” by a computational tool. Here we called this process as “knowledge extraction”. Once again, we borrow the experience in the medical field where methods to extract knowledge from textual data have been developed in the past decade. The missing element for us is the availability of a complete dictionary of NPP related concepts (in addition to the MBSE models presented earlier) that can put “text into context”. In FY23, such dictionary has been developed along with all the computational elements required for knowledge extraction. Lastly, once numeric and textual ER data elements have been processed and “understood”, then the last step is the discovery of possible cause-effect relations among them. This is performed by observing if a logical connection through the MBSE models exists, and if the

97 MATHEMATICS AND COMPUTING↗