Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parsing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Procedure Parsing: A Method for Parsing Handwritten Documents into Computer-Based Procedures

The nuclear industry is heavily procedure driven, where almost everything has a step-by-step instruction that is expected to be followed in detail. Historically, these procedures were printed on paper copies. Recently, the industry transitioned towards electronic copies (i.e., PDFs on tablets). One major drive for this transition is the introduction of human error and loss of situation awareness when using paper copies. However, electronic copies of documents inherently have the same error traps as their paper cousins. Therefore, there is an increased interest in a way to utilize the information in the step-by-step guidance, but to present it in a dynamic manner that guides the user and adapts to any encountered conditions. Researchers at Idaho National Laboratory propose a flexible, automated method based on document parsing and augmented by natural language processing (NLP) techniques, to address these shortcomings and capitalize on these recent advancements in machine learning. The proposed method provides a cost-effective solution for computer-assisted procedure parsing of hand-written control room procedures, originally authored in Word or PDF formats, into instructions that can be displayed as computer-based procedures (CBP) in a modern graphical user interface. The researchers devised, implemented and demonstrated the Operating Procedure Extender for Novel Systems (OPENS) method in 2020. The key to OPENS is to map the original procedure text into a context-free grammar, tying content to equipment, locations, and other steps, actions, etc. This formal grammar is then used to isolate and define keywords and actions verbs, such as “measure” or “evaluate” and tie them to specific equipment referenced within that step or located in other steps, substeps, actions, subactions and tables throughout the procedure. OPENS generates an abstract syntax tree from the document which it uses to store a copy of this information in the open-standard, machine-readable and human-readable file formats XML and JSON. The XML is useful to preserve the relational aspects of the procedure for referencing tables and branching information so the user can be directed to the next appropriate active step based on the values entered for that step and previous steps. The JSON is useful for storing and exchanging data objects used to track responses to previous steps and state changes in simulated environments. In future iterations, these formats can also be used for storing more detailed information about input during plant operation or simulation. The techniques the researcher developed could further be improved by integration of recent advancements in machine learning. NLP methods could standardize documents, correct for grammatical error, and provide automated semantic validation. The researcher expects that self-supervised techniques applied to collections of natural language instructions could strengthen the model with broader context. All these methods together give us a practical way to automatically extract protocols from documents and user interactions, empowering researchers, procedure writers and nuclear operators while moving the industry forward.

99 GENERAL AND MISCELLANEOUS↗

A Practical guide to Parsing MCNP Inputs: Lessons Learned from Implementing Context-Free Parsing in MontePy

Monte Carlo N-Particle (MCNP) is a widely used Monte Carlo transport solver that began development in the 1960’s. Due to this MCNP input files uses a custom input syntax, for which there are no off-the-shelf parsing libraries available. For MontePy to create an effective Object-Oriented interface for MCNP input files, an context-free parser was implemented to be able to fully parse the files. MontePy uses a number of shortcuts and optimizations to avoid creating a single universal input file parser. . These lessons can be applied to working with the many other custom input syntax languages persistent throughout the nuclear industry.

97 MATHEMATICS AND COMPUTING↗

Easy_PERT: a Python tool for writing PERT cards and parsing PERT card results [Slides]

This presentation begins by providing an overview of the PERT card. The PERT card uses differential operator method to compute first- and second-order tally variations due to density, composition, and reaction cross-sections. It is possible to have multiple PERT cards in one MCNP input deck to study tally variations for several sets of nuclides, reactions, and energy ranges. Furthermore, the METHOD option tells MCNP to calculate either the perturbed tally (METHOD=-1, -2, -3) or the change in the unperturbed tally (METHOD=1, 2, 3). In summation, a powerful use-case for the MCNP code PERT card is that it facilitates calculating tally sensitivities to nuclear data. Writing PERT card entries and parsing output MCTAL files is tedious and error prone. however, Easy_PERT makes use of existing tools (Faust and MCNPTools) to handle writing PERT card entries and parsing the output MCTAL files. The PERT card is early in the development process and planned upcoming capabilities include calculating sensitivities and combining MCTAL files from separate runs into one JSON file.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Complex Parsing for In-Network Acceleration of High-Energy Physics Experiments

This paper describes a novel application and evaluation of programmable networking in High-Energy Physics (HEP): a complete parser for the custom packet format used by Fermilab’s DUNE experiment. Notably, this parser is implemented on a Tofino programmable network switch and evaluated on the FABRIC testbed by using network traffic generated by the ICEBERG DUNE prototype. The parsed network traffic consists of Jumbo Ethernet frames that contain digitizations of sensor readings from ICEBERG’s detector.This work is an early investigation into providing in-network processing support for HEP experiments. The paper describes DUNE’s custom packet format, the challenges encountered when implementing a parser for that format, and an exploration of the techniques that are needed to overcome those challenges. We identify performance bottlenecks and discuss directions for future research.

Sagstad, Bjoern [IIT, Chicago] (ORCID:000900033610↗

Auto Procedure Parsing: A Natural Language Processing Approach

Nuclear Power Plant (NPP) operating procedure is “a set of rules that describes how actions on the plant should be made if a certain system goal should be accomplished. U.S. NPPs use paper-based procedures (PBPs). PBPs are difficult to use. Common errors with PBPs are: following the wrong procedure, omit a step etc. Computer based procedures (CPBs) offer great improvement in ensuring plant safety. Some studies rely on experienced operators to understand the procedure content and then reorganize the procedure with digitally executable capabilities. Other studies utilize the procedure format to design rules to extract information from the operating procedures. Existing studies in procedure parsing are manual, laborious and extract limited information. This study aims to automatically extract critical information from operating procedures for generating computer interpretable representation of procedures. Such representation can further be used for automatic dynamic human reliability analysis and CPB design etc.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

TopTemp: Parsing Precipitate Structure from Temper Topology

Technological advances are in part enabled by the development of novel manufacturing processes that give rise to new materials or material property improvements. Development and evaluation of new manufacturing methodologies is labor-, time-, and resource-intensive expensive due to complex, poorly defined relationships between advanced manufacturing process parameters and the resulting microstructures. In this work, we present a topological representation of temper (heat-treatment) dependent material micro-structure, as captured by scanning electron microscopy, called TopTemp. We show that this topological representation is able to support temper classification of microstructures in a data limited setting, generalizes well to previously unseen samples, is robust to image perturbations, and captures domain interpretable features. The presented work outperforms conventional deep learning baselines and is a first step towards improving understanding of process parameters and resulting material properties.

Kassab, Lara↗

EV-ELM (Electric Vehicle Policies with the Energy Language Model) [SWR-25-156]

Electric Vehicle Policies with the Energy Language Model (EV-ELM) leverages previous work using Large Language Models (LLMs) to find, download, and parse policy information related to energy infrastructure. In this application, we use LLMs to find policy documents related to the permitting and installation of electric vehicle charging infrastructure. This software contains the code to find, download, and parse these documents, while a related data record in the Open Energy Data Initiative (OEDI) will include the resulting output dataset that can be used for downstream analysis. The EV-ELM repository contains code for the EV-ELM project, which focuses on retrieving and processing EV permitting processes using large language models. The project is composed of two pipelines: (1) a web scraping pipeline for discovering and downloading EV permitting documents, and (2) a document parsing and extraction pipeline that processes the downloaded files to produce structured data. The web scraping pipeline is designed to extract relevant information from various websites, while the document parsing pipeline processes and analyzes the extracted documents to derive meaningful insights. Both pipelines depend on the NLR elm repository, which provides essential tools and functionalities for handling and processing the data. The web scraping pipeline is a modified version of the ordinance_gpt example within the elm repository. It has been adapted to fit the specific requirements of the EV-ELM project, ensuring that it effectively captures and processes the necessary information related to EV permitting.

Olson, Reid [National Laboratory of the Rockies (N↗

AdaParse

SF-25-127AdaParse (Adaptive Parallel PDF Parsing and Resource Scaling Engine) enables scalable, high-accuracy PDF parsing. AdaParse is a data-driven strategy that assigns an appropriate parser to each document, offering high accuracy for any computational budget. Moreover, it offers a workflow of various PDF parsing software that includes extraction tools: PyMuPDF, pypdf traditional OCR: Tesseract, modern OCR (e.g., Vision Transformers): Nougat and Marker

Siebenschuh, Carlo↗

A decay database of coincident γ–γ and γ–X -ray branching ratios for in-field spectroscopy applications

Current fieldable spectroscopy techniques often use single detector systems heavily impacted by interferences from intense background radiation fields. These effects result in low-confidence measurements that can lead to misinterpretation of the collected spectrum. To help improve interpretation of the fission products and short-lived radionuclides produced in a composite sample, a coincidence-database is being developed in support of a robust portable and X-ray coincidence detector system concurrently under development at the Pacific Northwest National Laboratory for in-field deployment. Hitherto, no database exists containing coincident γ–γ and γ–X-ray branching-ratio intensities on an absolute scale that will greatly enhance isotopic identification for in-field applications. As part of this project, software has been developed to parse all radioactive-decay data sets from the Evaluated Nuclear Structure Data File (ENSDF) archive to enable translation into a more useful JavaScript Object Notation (JSON) formats that more readily supports query-based data manipulation. The coincident database described in this work is the first of its kind and contains coincidence γ–γ and γ–X-ray intensities and their corresponding uncertainties, together with auxiliary metadata associated with each decay data set. The new JSON format provides a convenient and portable means of data storage that can be imported into analysis frameworks with relatively low overhead allowing for meaningful comparison with measured data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A Flang Plugin for Fortran Feature Characterization

As new compute systems are developed, there is still a need to compile and execute codes authored in Fortran on these leading edge systems. In order to achieve this, development of compilers that support the latest hardware is continuously under development. Though the specification of Fortran is extensive, it is helpful to compiler authors to be able to prioritize the development of key features in order to get certain codes deemed important, e.g., applications of interest to leadership computing facilities, executable on leading edge compute systems. Identifying key features though is largely done through querying software experts or users of the Fortran applications of interest, who then manually report what features are and are not present. This exercise can both time consuming and error prone. To automate this process, we present a compiler plugin to Flang, the Fortran frontend for LLVM. This plugin is a tool that operates on the parse tree representation generated by Flang and detects key features based on walking parse tree nodes that correspond to features of interest. We show the result of our tool on four applications, three of which were manually profiled by software experts. We show the discrepancies between our tool and the manual characterization of the three applications, as well as generate a characterization for an application not yet profiled. We intend to open-source our tool in order to invite the community to benefit from the tool and make contributions for other features.

Cabrera, Anthony [ORNL]↗

High-Throughput Computing: Case Study of Medical Image Processing Applications

HPC is designed for large-scale simulations using monolithic codes of tightly coupled processes highly optimized to deliver decreased time to solution. Medical image processing is not a traditional field of HPC. Similar to AI applications, medical image processing parses large datasets, typically multiple times, to support a variety of studies for classification, diagnosis or monitoring purposes. The convergence of AI, HPC and Big Data encouraged more fields using image processing to transition to HPC. However, not all applications benefit from the same optimizations. In this paper we focus on high throughput medical image processing applications that analyze a huge dataset of small MRI images and that require HPC systems to decrease the time of parsing the entire dataset and not individual MRIs. We show in this research the performance of running SLANT, an image processing application for a whole brain segmentation, on large-scale systems and highlight performance limitations. We present optimizations prioritizing throughput that exhibit a 3.5x speed-up on the Summit Supercomputer that can be used as a baseline for building a high-throughput execution framework for other HPC systems.

Predescu, Maria↗

AI-Batt (Autonomous Identification of Battery Life Models) [SWR 21-36]

Autonomous Identification of Battery Life Models (AI-Batt) AI-Batt is a MATLAB code base for developing lifetime models for batteries from accelerated aging data. The code base provides many functions for processing, visualizing, and modeling battery aging data, making the data processing, exploration, and modeling workflow substantially faster. These tools are tailored for working with battery aging data sets, which usually consist of many separate time-series for each cell, with many test conditions and possible replicates at each condition, which makes it difficult to simply process or visualize the data set. Complex modeling tasks, such as cross-validation, sensitivity analysis, and uncertainty quantification have been implemented to enable thorough statistical investigation of model predictions. Additionally, several machine-learning algorithms are implemented to autonomously identify suitable models via symbolic regression. Data processing functions automatically cast data from the struct data type, which is commonly used to store experimental data, but is not an acceptable input for most algorithms, to the table data type, which can be easily used as input to any optimization algorithm. Also, the data can be separated into time-invariant and time-variant data tables, which is helpful for exploring the data set as well as developing separate models for time-variant and time-invariant aging mechanisms. For example, in aging tests with constant temperature, temperature is a time-invariant experimental condition. Visualization tools enable plotting of data, model fits, and model simulations possible with single-line function calls, empowering data exploration of complex data sets with both time-varying and time-invariant trends. Plots can be automatically generated for the whole data set, or separated by data group (groups of test replicates) or individual data series. Data points or data series can be automatically colored by the value of a variable with a variety of color maps, and model predictions can also be colored by the value of a fit statistic. Comparisons between data sets and the predictions/simulations of different models on the same data set can be easily plotted as well. Distributions of parameter values from bootstrap resampling can be plotted to visualize the reliability of parameter estimation, or determine any correlations between parameters. Modeling tools handle the complex task of creating and parsing symbolic equations for modeling battery lifetime. Equations are parsed to grab relevant data variables, parameter values, or specified sub-models for input into optimization, evaluation, or simulation functions. Models can be optimized locally (one set of parameters for each data series), bi-level (some parameters shared across the data set), or globally (single set of parameters for all data). Functions implementing symbolic regression algorithms help users to discover effective model equations, even in poorly sampled, high-dimensional data.

Smith, Kandler [National Renewable Energy Lab. (NR↗

JGI-Trichoderma v1.0

There is a series of Python and bash scripts to parse genomics datasets used to evaluate the coevolution of gene families and the feature importance of gene families using an SVM classifier. - Cover analysis: takes a list of single-copy genes in a set of genomes, aligns and builds the gene trees to determine if two gene families have a signature of covariation with one another. It parses the files to run phykit cover script described here: https://jlsteenwyk.com/PhyKIT/usage/index.html - SVM-classifier: This Python script is an SVM-based genomic classifier designed for biological data analysis. It combines machine learning with feature selection to identify important genomic markers and classify biological samples. Core Functionality: The script uses Support Vector Machines from scikit-learn to classify genomic data, incorporating SelectKBest for automated feature selection and leave-one-out cross-validation for performance assessment. It operates in multiple modes: feature ranking, optimal combination discovery, and sample prediction. Primary Applications: Genomic sample classification and biomarker discovery Feature importance analysis in high-dimensional biological datasets Prediction of sample categories based on genomic profiles Research applications requiring robust classification of biological data Key Advantages: High-dimensional handling: SVMs excel with genomic data's typical high feature-to-sample ratios Integrated feature selection: Reduces noise and computational overhead while identifying key markers Probability estimation: Provides confidence scores essential for biological interpretation Validation robustness: Leave-one-out cross-validation ensures reliable performance metrics Operational flexibility: Multiple analysis modes support different research phases from exploration to prediction

Stecca Steindorff, Andrei [Lawrence Berkeley Natio↗

Zero Resistance Ammetry (ZRA) Measurements of Sediment Electrochemical Gradients, Old Woman Creek, Ohio, USA, June–November 2022

This dataset contains raw and processed zero resistance ammetry (ZRA) measurements collected from wetland sediments at Old Woman Creek, a freshwater estuary on Lake Erie, Ohio, USA, between June and November 2022. Measurements were obtained using a vertically deployed electrode array positioned at multiple depths within the sediment profile to capture electrochemical gradients associated with microbial activity and sediment geochemistry. The raw dataset consists of parsed instrument log files containing timestamps, electrode pair identifiers, and measured electrical potential (mV). The processed dataset includes standardized and quality-controlled values with instrument saturation limits removed and timestamps converted to ISO 8601 format. Electrode line identifiers were mapped to physical depths, enabling interpretation of depth-resolved electrochemical gradients. Instrument saturation values (−2048, −2047, 2047, and 2048 mV) were identified as measurement limits and excluded from quantitative analyses. All data parsing, processing, and quality control steps are documented in an accompanying R Markdown script, ensuring full reproducibility from raw instrument logs to final datasets.

EARTH SCIENCE > AGRICULTURE > SOILS > ELECTRICAL C↗

High-Throughput Custom Monitoring for the Mu2e TDAQ System

In this project we are studying the application of programmable network hardware to provide a custom monitoring capability for the Mu2e Trigger and Data Acquisition System (TDAQ) system. The goal of the Mu2e experiment is to search for a charged-lepton flavor violating processes where a negative muon converts into an electron in the field of an aluminum nucleus. This experiment is intended to improve by four orders of magnitude the search sensitivity reached so far. We have a working prototype of a system that provides high-throughput, custom monitoring for the Mu2e TDAQ system. The custom Mu2e network packet header format is parsed as it crosses the network switch. Parsing extracts bits that convey information about error states at read-out controllers (ROCs). This information is periodically relayed to the switch controller, which in turn alerts experiment operators.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

High-Throughput Custom Monitoring for the Mu2e TDAQ System

In this project we are studying the application of programmable network hardware to provide a custom monitoring capability for the Mu2e Trigger and Data Acquisition System (TDAQ) system. The goal of the Mu2e experiment is to search for a charged-lepton flavor violating processes where a negative muon converts into an electron in the field of an aluminum nucleus. This experiment is intended to improve by four orders of magnitude the search sensitivity reached so far. We have a working prototype of a system that provides high-throughput, custom monitoring for the Mu2e TDAQ system. The custom Mu2e network packet header format is parsed as it crosses the network switch. Parsing extracts bits that convey information about error states at read-out controllers (ROCs). This information is periodically relayed to the switch controller, which in turn alerts experiment operators.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

RE-INTEGRATE EMT Simulation Tool: Input Data Processing Layer for Bulk Power System

This paper introduces an advanced input data processing layer for EMT simulations of large-scale bulk power systems. The paper proposes two versions of the RE-INTEGRATE EMT simulation tool, RE-INTEGRATE Gen-0 and RE-INTEGRATE Gen-1, which are developed to enhance simulation generalizability, scalability, and accuracy. The framework leverages a generic class design for components to incorporate linear equations, which are generated by discretizing the Differential-Algebraic Equations (DAEs) that represent the dynamics of the components. In addition, the framework employs a parsing algorithm that parses a power system’s raw and dyr files to generate a connectivity graph which is then traversed to form the overall system’s dynamics. The proposed input data processing layer is used to simulate the IEEE 39-bus test system. The obtained results demonstrate the framework’s capability to achieve simulation scalability and accuracy. Further, the results indicate that EMT simulations performed using the proposed automations can effectively handle complex grid configurations.

Mishra, Rahul [ORNL] (ORCID:0000000328205932)↗