Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Carbon Management Projects (CONNECT) Database and Explorer

Overview The Carbon Management Projects (CONNECT) Toolkit is an online exploratory visualization tool developed by the U.S. Department of Energy's (DOE) Office of Fossil Energy and Carbon Management (FECM) with support from other federal agencies such as the U.S. Environmental Protection Agency (EPA) and the U.S. Department of Transportation (DOT). It provides a single point of access to authoritative information on federal agency investment in a portfolio of research, development, and demonstration (RD&D) projects that have been publicly announced to advance technologies for point source carbon capture, carbon dioxide removal, transport, storage, and conversion, collectively referred to as carbon management. The RD&D programs covered in this tool are authorized by annual congressional appropriations ("Base Program") and the 2021 Infrastructure Investment and Jobs Act (IIJA). The tool also incorporates public information on other federal initiatives, such as the Regional Clean Hydrogen Hubs, and public information released by other government agencies, such as the Environmental Protection Agency's (EPA) and Primacy States’ Underground Injection Control Class VI permits and EPA’s facility level greenhouse gas (GHG) emissions. Developed in a geographic information system, the tool organizes carbon management projects into five groups based on the primary technology that a project aims to advance, each visually represented as a digital layer ("carbon management project layer"). Only federally funded projects are included, which can be awarded projects that are completed or ongoing, or projects that have been selected but are currently under negotiation. Project information can be viewed in the map or in the attribute table below it when turned on. In the map view, each project is displayed at either its host site (for field work), where available, or its performer site (project lead's location, further explained in the table below). Host sites and performer sites are represented in distinct icons. Several reference layers offer additional public information on infrastructural and natural resource environment for carbon management. These reference layers, combined with multiple geographical basemaps, enable users to visualize the carbon management project layers in context. Carbon management project information will be updated monthly based on feedback and information availability. Carbon management project layers Point Source Carbon Capture (PSC) This layer contains DOE-funded projects focused on capturing carbon dioxide (CO2) from power plants or industrial facilities. Carbon Dioxide Removal (CDR) This layer contains DOE-funded projects focused on capturing CO2 from the atmosphere, including direct air capture (DAC) and DAC hubs, direct ocean capture, enhanced mineralization, and biomass carbon removal and storage. For projects with multiple host sites, each of the sites are displayed individually with the project cost and cost sharing information representing the total for the entire project. Carbon Transport This layer contains DOE- and DOT-funded projects focused on CO2 transport. The Transport Research and Development sublayer contains projects that do not involve physical infrastructure; the Proposed Transport Corridor sublayer contains projects for which either a route for the transport infrastructure has been proposed or a general area for the transport infrastructure has been identified. Carbon Storage This layer contains DOE-funded key projects focused on CO2 storage. For projects with multiple field-work sites, each of the sites are displayed individually on the map with the project cost and cost sharing information representing the overall total for the entire project. Carbon Conversion This layer contains DOE-funded projects focused on converting CO2 into economically valuable products. Reference layers The following layers provide additional information in the geographic proximity of carbon management projects. Users should reference the original sources for more details (weblinks provided below and in pop-up windows on the map). Regional Clean Hydrogen Hub and Facility These layers illustrate the approximate areas of the Regional Clean Hydrogen Hubs announced by DOE's Office of Clean Energy Demonstrations (OCED) and the approximate locations of individual facilities that constitute the hubs (see "Where are the H2Hubs located?" on the webpage linked above). EPA Facility Level GHG Emissions (direct emitter) This layer shows direct CO2 emissions from stationary sources in 2022, using data extracted from EPA's Facility Level Information on GreenHouse gases Tool (FLIGHT). Captured and injected CO2 are not deducted from direct emitters’ total emissions. Contact EPA for additional details. Underground Injection Control Class VI permit/permit application This layer shows the locations of CO2 injection wells that are granted or in the process of applying for an Underground Injection Control Class VI permit by EPA or a Primacy State (currently Louisiana, North Dakota, and Wyoming). The URLs for the permits or permit applications are provided in the pop-up windows associated with the well locations. Contact EPA for additional details. Carbon Storage Resource This layer contains information on prospective CO2 storage resources in saline formations and oil and gas reservoirs provided by the National Carbon Sequestration Database and Geographic Information System (NATCARB) spatial database. Contact NETL for additional details. Existing CO2 pipeline This layer shows active CO2 pipelines based on information digitized from the map issued by the Pipeline and Hazardous Materials Safety Administration (PHMSA). Contact PHMSA for additional details.

Carbon Conversion↗

Laser Confocal Microscopy Uncertainty Quantification Study

At Los Alamos National Laboratory (LANL), the Storage Safety and Engineering (SSE) team completes annual surveillance on a subset of in-use interim nuclear material storage containers in fulfilment of requirements outlined in DOE Manual M 441.1-1. The containers are selected through several methods, such as subject matter expert judgement, random selection, and trending items. Following these selections, the SSE team has the capacity to complete surveillance on 15-20 containers each fiscal year, composed of a combination of SAVY-4000 and Hagan storage containers. Through previous work, the stainless-steel components of the containers have been identified as life limiting components, with an emphasis on the thin-walled bodies. The team is focused on understanding the extent of general and pitting corrosion, due to observations of extensive corrosion from stored contents and bag-out-bag degradation. Quantifying corrosion effects on the thin-walled stainless steel container bodies, and understanding potential impacts to the respective design release rates and design qualification release rates is paramount to the team. To date, destructive examination (DE) has proven to be the most insightful method for developing an understanding on the extent of corrosion on used containers. To standardize this process, the SSE team developed a destructive examination guide for analyzing stainless steel components of the containers. Corroded containers of interest are identified during surveillance activities and set aside for sectioning and characterization. Following sectioning, a major step in the DE workflow is the utilization of laser confocal microscopy for scanning corroded samples of interest and extracting data on pits, such as count, depth, and equivalent diameter. Adhering to the techniques outlined in the DE guide, analysis has been completed on two Hagans and one SAVY-4000 container, with the maximum pit depth recorded as 139.1 ± 22.82 μm on a 17.5 year old Hagan. The findings from the completed destructive examinations will be utilized to support lifetime extension efforts of the SAVY-4000 as the team can better estimate corrosion rates and effects over time based on stored contents and age. Due to the implications of observing extreme pit depths that approach the nominal container body thickness of .0299 inches (0.759 mm) or minimum container thickness of 0.236” (0.6 mm), high confidence in the LCM measurements is desired. Through testing outlined in, it was concluded that the total error ascribed to the 20x objective when conducting large image mapping on the Keyence VK-X3050 laser confocal microscope (LCM) relative to a 50x objective (reference) is 16.4% (± 8.73%). For shallow features on the order of pristine SAVY surface defects (i.e. 5 μm), this uncertainty is appropriate. However, this conservative estimate of total error poses a fundamental concern for pit depths that approach the thickness of the measured samples. That is, with the measurement uncertainty currently employed on all measurements, the LCM would be unable to resolve if a pit with a depth of 515 μm is through wall. Standard step height samples were procured and used in the present study to assess the resolution and repeatability of height measurements. Understanding the resolution and repeatability of height measurements was the first focus of the team as it relates directly to pit depth, which is of primary concern. Calibration gratings were procured to evaluate the resolution and repeatability of measurements in the X and Y axes of the LCM stage. The results of the depth uncertainty study were conducted first and presented in the subsequent sections. The planar uncertainty study is appended to the depth study with conclusions from both summarized at the end of the report.

36 MATERIALS SCIENCE↗

Measurement of $\nu_\mu/\bar\nu_\mu$ CC double-differential cross sections on MINERvA hydrocarbon target for the shallow inelastic scattering background region

Cross section measurements are essential for all neutrino oscillation experiments. In fact, uncertaintiesassociated to cross section model parameters constitute one of the dominant sources oferrors in current oscillation analyses. In particular, understanding neutrino-induced pion productionin the kinematic regime known as shallow inelastic scattering (SIS) is critical for improvingneutrino interaction modeling in event generators. In this study, 416,233 (237,468) muon neutrino(antineutrino) interactions are measured in a SIS background region, predominantly made ofbaryon resonances. The analyzed datasets were collected from 2013 to 2019, comprising neutrinosgenerated by the Fermilab NuMI facility, with mean energy of 6 GeV, and the interactions occurredon the MINERvA hydrocarbon target. The measurements are presented as double-differential crosssections in terms of the outgoing muon longitudinal and transverse momentum components, aswell as the Bjorken x and y variables. Comparisons between the extracted data and predictionsfrom several generators reveal significant discrepancies across most kinematic bins.

Souza Correia, Souza Correia, Daniel [Rio de Janei↗

HTESP (High-throughput electronic structure package): A package for high-throughput ab initio calculations

High-throughput ab initio calculations are the indispensable parts of data-driven discovery of new materials with desirable properties, as reflected in the establishment of several online material databases. The accumulation of extensive theoretical data through computations enables data-driven discovery by constructing machine learning and artificial intelligence models to predict novel compounds and forecast their properties. Efficient usage and extraction of data from these existing online material databases can accelerate the next stage materials discovery that targets different and more advanced properties, such as electron–phonon coupling for phonon-mediated superconductivity. However, extracting data from these databases, generating tailored input files for different ab initio calculations, performing such calculations, and analyzing new results can be demanding tasks. Here, in this work, we introduce a software package named “HTESP” (High-Throughput Electronic Structure Package) written in Python and Bash languages, which automates the entire workflow including data extraction, input file generation, calculation submission, result collection and plotting. Our HTESP will help speed up future computational materials discovery processes.

36 MATERIALS SCIENCE↗

Performance evaluation of automated data-driven feature extraction and selection methods for practical and scalable building energy consumption prediction models

Here, this study quantifies the impact of automated feature engineering methods (feature extraction and selection) on the quality and accuracy of machine learning models that predict building energy consumption. The case study compares model performance for three main scenarios: baseline (no feature extraction and selection), feature extraction only, and feature extraction combined with feature selection (filter and/or wrapper methods) for fully trained machine learning models for 200 metered/sub-metered energy measurements across 118 real buildings. For consistency, the same machine learning model architecture (a black box deep learning neural network with probabilistic forecast output) was used for all scenarios. Based on results, all feature engineering methods provided noticeable prediction accuracy improvements (e.g., 29%-68% median prediction improvement) compared to baseline scenarios. However, in this application, feature selection methods provide little practical value due to their limited performance gains and high computational cost. Smarter algorithm development supported by better computational environments will be needed before feature selection methods can reliably and efficiently improve predictive model performance.

97 MATHEMATICS AND COMPUTING↗

Extraction of Vibration Data with Imaging

To date, the primary sensing technology used to measure the vibration response has been accelerometers and strain gages mounted directly to the structure and using either wired or, more recently, wireless telemetry. Cost issues with these sensors and the associated data acquisition systems typically limit the numbers that are deployed on in situ structures. Although there are a few structures with larger sensing counts that in some cases exceed over 1000 sensors, more typical numbers range from ten to one hundred sensors resulting in low spatial resolution when they are applied to physically large systems. When one considers that nuclear power plant structures usually have complex geometries, material properties, connectivity and boundary conditions, it is clear these current approaches to vibration measurements can only provide limited information about a system’s dynamics response characteristics. As an alternative, many non-contact measurement technologies have emerged, including point wise measurement methods such as Global Positioning System (GPS), microwave interferometry, and laser Doppler vibrometry (LDV), as well as simultaneous full-field measurement methods such as electronic speckle pattern interferometry, holography interferometry, and muon tomography, some of which can provide high spatial resolution measurements. Among these methods, digital video imaging techniques have emerged as a feasible solution for full-field vibration measurements that provide significantly more detailed dynamic response information because every pixel becomes a measurement point. Furthermore, recent advances in image processing and computer vision algorithms have been successfully used to process video data for experimental and operational modal analysis. Such full-field measurements have the potential to significantly improve many current structural assessment procedures including system identification (modal parameter estimation), structural health monitoring, load reconstruction, model validation, and model updating. Furthermore, more recent full-field imaging techniques can be accomplished with relatively low-cost, commercially-available off-the-shelf cameras. However, these measurement procedures have other limitations that must be considered such as the ability to only measure visibly accessible points on a structure and a more limited dynamic range and bandwidth than can be achieved with accelerometers or strain gages.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Multioutput Convolutional Neural Network for Improved Parameter Extraction in Time-Resolved Electrostatic Force Microscopy Data

Time-resolved scanning probe microscopy methods, like time-resolved electrostatic force microscopy (trEFM), enable imaging of dynamic processes ranging from ion motion in batteries to electronic dynamics in microstructured thin film semiconductors for solar cells. Reconstructing the underlying physical dynamics from these techniques can be challenging due to the interplay of cantilever physics with the actual transient kinetics of interest in the resulting signal. Previously, quantitative trEFM used empirical calibration of the cantilever or feed-forward neural networks trained on simulated data to extract the physical dynamics of interest. Both these approaches are limited by interpreting the underlying signal as a single exponential function, which serves as an approximation but does not adequately reflect many realistic systems. Here, we present a multi-branched, multi-output convolutional neural network (CNN) that uses the trEFM signal in addition to the physical cantilever parameters as input. The trained CNN accurately extracts parameters describing both single-exponential and bi-exponential underlying functions, and more accurately reconstructs real experimental data in the presence of noise. This article demonstrates an application of physics-informed machine learning to complex signal processing tasks, enabling more efficient and accurate analysis of trEFM.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

To Derive or Not to Derive: I/O Libraries Take Charge of Derived Quantities Computation

The ever-increasing volume of data produced by HPC simulations necessitates scalable methods for data exploration and knowledge extraction. Scientific data analysis often involves complex queries across distributed datasets, requiring manipulation of multiple primary variables and generating derived data that needs to be handled efficiently, creating challenges for applications that need to parse many large datasets. Relying on individual applications to handle all intermediate data generally leads to redundant computations across studies and unnecessary data transfers. In this paper, we investigate the performance of different approaches where applications define derived variables as quantities of interest (QoIs) and offload the computation and transfer of these QoIs to the I/O library. This significantly reduces redundancy and optimizes data movement across the distributed storage and processing infrastructure by allowing control over when and where derived variables are computed. We present a detailed analysis of the performance-storage trade-offs associated with different solutions and showcase results for our study on two large-scale datasets created from climate and combustion simulations.

Gainaru, Ana↗

A Knowledge Graph Approach to Analyze Systems and Assets Health

Nuclear power plants collect large amounts of equipment reliability data elements that contain information on the statuses of component, assets, and systems. All these data elements precisely record asset and system performance and health throughout the lifecycle of those assets and systems. However, several challenges have proved to be roadblocks to this process. While some of these challenges are technical in nature (i.e., data are often distributed over several physical servers or databases), others are conceptual in nature (i.e., data elements come in different formats, numeric or textual), and measured values have different scales (e.g., vibration spectra and oil temperature). This paper directly focuses on the integration of numeric and textual data elements in order to assist plant system engineers in analyzing equipment reliability data. This task begins with preprocessing the data by extracting knowledge from textual data via natural language processing methods and quantifying system, asset, and component health based on numeric data. We then employed model-based system engineering (MBSE) models of systems and assets to identify their architecture and functional (i.e., cause and effect) relations. Data elements were then associated with a single MBSE graph element, based on their nature. This bonding of MBSE models and data elements constitutes a first-of-its-kind knowledge graph of a nuclear power plants system, with data elements being organized in a structured manner that enables system engineers to identify cause-effect trends in data elements and carry out appropriate actions in response.

97 - MATHEMATICS AND COMPUTING↗

A Model Based Approach to Extract Health Information from Textual Data

In current nuclear power plants (NPPs) a large amount of condition-based data is being generated and stored to assess and monitor component health and performance. The format of this data can be either numeric (e.g., pump vibration data) or textual (e.g., condition report which assess component health). While assessing component health from numeric data can be performed with a large variety of methods, the extraction of information from textual data still remains a challenge. Natural language processing (NLP) methods are starting to be deployed in current NPPs mainly to filter out incident reports (IRs) that are not safety related by employing supervised machine learning methods. However, these methods do not really provide the quantitative information that might be contained in IRs. This paper presents an approach to extract information from textual data (e.g., from IRs, maintenance reports) that is based on NLP data analytics methods coupled with model-based system engineer (MBSE) models. NLP methods are employed to perform syntactic and semantic analyses. Syntactic analysis analyzes the grammatical structure of a sentence; such analysis includes: part of speech (POS) tagging (i.e., identification of grammatic elements of each string - e.g., nouns, verbs), named entity recognition (i.e., identification of text entities - e.g., names, dates, events), and relation extraction (e.g., coreference resolution). On the other hand, semantic analysis is designed to analyze the logic structure of a sentence. Through a specific set of rules, our methods can identify whether a sentence contains health information of a component (e.g., degraded performance, anomaly behavior) or the causal relationship between two events (i.e., a cause-effect pair). An innovative element of our approach is that semantic analysis relies on MBSE models to identify links between textual elements. MBSE are diagrams designed to represent system and component dependencies (from both a form and functional point of view). In our approach, MBSE models emulate system engineer knowledge about component/system architecture. This paper presents in detail how the integration of NLP methods and MBSE models is performed. Few analysis examples focusing on centrifugal pumps are presented.

97 - MATHEMATICS AND COMPUTING↗

Stoichiometrically-informed symbolic regression for extracting chemical reaction mechanisms from data

A data-driven computational method is introduced to extract chemical reaction mechanisms from time series chemical concentration data. It is realized through the use of dynamic symbolic regression in which a sparse analytical form for a dynamical system is discoverable from the underlying data. We specifically develop the stoichiometrically-informed symbolic regression (SISR) method to address a standing challenge in complex chemical reaction networks: given a time-series dataset of concentrations of several components, what is the mechanism and the associated rate constants? SISR finds the optimal mechanism, kinetic equations and rate constants by combining differential optimization with a genetic optimization approach that searches a symbolic space of possible reaction mechanisms. Use of SISR in several paradigmatic examples spanning linear and nonlinear reaction schemes results in excellent agreement between true and predicted mechanisms, including when the method is applied to noisy data. The advantages of a stoichiometrically-informed approach such as SISR to address reaction discovery is illustrated through comparison with the use of generic state-of-the-art data-driven approaches.

36 MATERIALS SCIENCE↗

LogPath: Log data based energy consumption analysis enabling electric vehicle path optimization

Vehicle navigation and path optimization require a more meticulous approach when it deals with EVs (electric vehicles) and SDVs (software-defined vehicles), due to lengthy charging times and the lack of charging infrastructure. Long-distance freight EV trucking needs path guidance with accurate energy consumption estimates to prevent charging-related failures. We developed a novel energy consumption estimation approach that only uses battery log data to extract major vehicle parameters to increase EV navigation accuracy without additional sensors. This is enabled by extracting multiple drive modes from the log data for analysis. The system provides 1) routes, 2) charge locations, 3) charging times, and 4) optimal vehicle speeds that guarantee the shortest travel time. Here we successfully validated the system using log data collected from an EV and Tesla's Supercharging map in the US and compared it with the commercially available navigation system, Tesla's trip planner, whose capabilities solely include charging time and routing.

EV (Electric vehicles) navigation↗

From Text to Maps: LLM-Driven Extraction and Geotagging of Epidemiological Data

Epidemiological datasets are essential for public health analysis and decision-making, yet they remain scarce and often difficult to compile due to inconsistent data formats, language barriers, and evolving political boundaries. Traditional methods of creating such datasets involve extensive manual effort and are prone to errors in accurate location extraction. To address these challenges, we propose utilizing large language models (LLMs) to automate the extraction and geotagging of epidemiological data from textual documents. Our approach significantly reduces the manual effort required, limiting human intervention to validating a subset of records against text snippets and verifying the geotagging reasoning, as opposed to reviewing multiple entire documents manually to extract, clean, and geotag. Additionally, the LLMs identify information often overlooked by human annotators, further enhancing the dataset’s completeness. Our findings demonstrate that LLMs can be effectively used to semi-automate the extraction and geotagging of epidemiological data, offering several key advantages: (1) comprehensive information extraction with minimal risk of missing critical details; (2) minimal human intervention; (3) higher-resolution data with more precise geotagging; and (4) significantly reduced resource demands compared to traditional methods.

Harrod, Karly↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

PDF Entity Annotation Tool (PEAT)

While different text mining approaches – including the use of Artificial Intelligence (AI) and other machine based methods - continue to expand at a rapid pace, the tools used by researchers to create the labeled datasets required for training, modeling, and evaluation remain rudimentary. Labeled datasets contain the target attributes the machine is going to learn; for example, training an algorithm to delineate between images of a car or truck would generally require a set of images with a quantitative description of the underlying features of each vehicle type. Development of labeled textual data that can be used to build natural language machine learning models for scientific literature is not currently integrated into existing manual workflows used by domain experts. Published literature is rich with important information, such as different types of embedded text, plots, and tables that can all be used as inputs to train ML/natural language processing (NLP) models, when extracted and prepared in machine readable formats. Currently, both normalized data extraction of use to domain experts and extraction to support development of ML/NLP models are labor intensive and cumbersome manual processes. Automatic extraction of data and information from formats such as PDFs that are optimized for layout and human readability, not machine readability. The PDF (Portable Document Format) Entity Annotation Tool (PEAT) was developed with the goal of allowing users to annotate publications within their current print format, while also allowing those annotations to be captured in a machine-readable format. One of the main issues with traditional annotation tools is that they require transforming the PDF into plain text to facilitate the annotation process. While doing so lessens the technical challenges of annotating data, the user loses all structure and provenance that was inherent in the underlying PDF. Also, textual data extraction from PDFs can be an error prone process. Challenges include identifying sequential blocks of text and a multitude of document formats (multiple columns, font encodings, etc.). As a result of these challenges, using existing tools for development of NLP/ML models directly from PDFs is difficult because the generated outputs are not interoperable. We created a system that allows annotations to be completed on the original PDF document structure, with no plain text extraction. The result is an application that allows for easier and more accurate annotations. In addition, by including a feature that grants the user the ability to easily create a schema, we have developed a system that can be used to annotate text for different domain-centric schemas of relevance to subject matter experts. Different knowledge domains require distinct schemas and annotation tags to support machine learning.

97 MATHEMATICS AND COMPUTING↗

Simulation-trained machine learning models for Lorentz transmission electron microscopy

Understanding the collective behavior of complex spin textures, such as lattices of magnetic skyrmions, is of fundamental importance for exploring and controlling the emergent ordering of these spin textures and inducing phase transitions. It is also critical to understand the skyrmion–skyrmion interactions for applications such as magnetic skyrmion-enabled reservoir or neuromorphic computing. Magnetic skyrmion lattices can be studied using in situ Lorentz transmission electron microscopy (LTEM), but quantitative and statistically robust analysis of the skyrmion lattices from LTEM images can be difficult. In this work, we show that a convolutional neural network, trained on simulated data, can be applied to perform segmentation of spin textures and to extract quantitative data, such as spin texture size and location, from experimental LTEM images, which cannot be obtained manually. This includes quantitative information about skyrmion size, position, and shape, which can, in turn, be used to calculate skyrmion–skyrmion interactions and lattice ordering. We apply this approach to segmenting images of Néel skyrmion lattices so that we can accurately identify skyrmion size and deformation in both dense and sparse lattices. The model is trained using a large set of micromagnetic simulations as well as simulated LTEM images. This entirely open-source training pipeline can be applied to a wide variety of magnetic features and materials, enabling large-scale statistical studies of spin textures using LTEM.

McCray, Arthur R. C. (ORCID:0000000160774698)↗

Investigation of the Performance and Explainability Tradeoffs for Machine-Learning Models for Predictive Maintenance of Circulating Water Systems in Nuclear Power Plants

Predictive maintenance (PdM) has shown great potential for achieving substantial cost savings and enhancing the economic competitiveness of nuclear power plants (NPPs) in today's energy market. Among the different modeling approaches that exist, machine learning (ML) tools in particular have a demonstrated ability to handle high dimensional and multivariate data and to extract hidden relationships within data in industrial environments. While ML methods show great potential, their lack of explainability---especially for black-box models---is a major hurdle to their adoption. Moreover, considering the supposed trade-off between explainability and performance challenges, careful consideration must be made as to which of these quality aspects takes precedence in light of multiple modeling options, resource availability, and domain characteristics. The present work evaluates the performance of six ML models, each with a different degree of explainability, in classifying the conditions of circulating water pumps (CWPs) by utilizing sensor data from nuclear power plants. To determine the drivers behind the trade-offs presented by this array of models, this work also tests different combinations of CWP units as the training and testing data, degrees of data imbalance, and objective functions for hyperparameter tuning. It was found that black-box models tend to afford superior performance in cases where there are far more instances of one type of labeled data than of any other type. It is recommended that a guided procedure be followed for designing and delivering an ML system that is sufficiently explainable to all involved stakeholders.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗