Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parsing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Procedure Parsing: A Method for Parsing Handwritten Documents into Computer-Based Procedures

The nuclear industry is heavily procedure driven, where almost everything has a step-by-step instruction that is expected to be followed in detail. Historically, these procedures were printed on paper copies. Recently, the industry transitioned towards electronic copies (i.e., PDFs on tablets). One major drive for this transition is the introduction of human error and loss of situation awareness when using paper copies. However, electronic copies of documents inherently have the same error traps as their paper cousins. Therefore, there is an increased interest in a way to utilize the information in the step-by-step guidance, but to present it in a dynamic manner that guides the user and adapts to any encountered conditions. Researchers at Idaho National Laboratory propose a flexible, automated method based on document parsing and augmented by natural language processing (NLP) techniques, to address these shortcomings and capitalize on these recent advancements in machine learning. The proposed method provides a cost-effective solution for computer-assisted procedure parsing of hand-written control room procedures, originally authored in Word or PDF formats, into instructions that can be displayed as computer-based procedures (CBP) in a modern graphical user interface. The researchers devised, implemented and demonstrated the Operating Procedure Extender for Novel Systems (OPENS) method in 2020. The key to OPENS is to map the original procedure text into a context-free grammar, tying content to equipment, locations, and other steps, actions, etc. This formal grammar is then used to isolate and define keywords and actions verbs, such as “measure” or “evaluate” and tie them to specific equipment referenced within that step or located in other steps, substeps, actions, subactions and tables throughout the procedure. OPENS generates an abstract syntax tree from the document which it uses to store a copy of this information in the open-standard, machine-readable and human-readable file formats XML and JSON. The XML is useful to preserve the relational aspects of the procedure for referencing tables and branching information so the user can be directed to the next appropriate active step based on the values entered for that step and previous steps. The JSON is useful for storing and exchanging data objects used to track responses to previous steps and state changes in simulated environments. In future iterations, these formats can also be used for storing more detailed information about input during plant operation or simulation. The techniques the researcher developed could further be improved by integration of recent advancements in machine learning. NLP methods could standardize documents, correct for grammatical error, and provide automated semantic validation. The researcher expects that self-supervised techniques applied to collections of natural language instructions could strengthen the model with broader context. All these methods together give us a practical way to automatically extract protocols from documents and user interactions, empowering researchers, procedure writers and nuclear operators while moving the industry forward.

99 GENERAL AND MISCELLANEOUS↗

A Practical guide to Parsing MCNP Inputs: Lessons Learned from Implementing Context-Free Parsing in MontePy

Monte Carlo N-Particle (MCNP) is a widely used Monte Carlo transport solver that began development in the 1960’s. Due to this MCNP input files uses a custom input syntax, for which there are no off-the-shelf parsing libraries available. For MontePy to create an effective Object-Oriented interface for MCNP input files, an context-free parser was implemented to be able to fully parse the files. MontePy uses a number of shortcuts and optimizations to avoid creating a single universal input file parser. . These lessons can be applied to working with the many other custom input syntax languages persistent throughout the nuclear industry.

97 MATHEMATICS AND COMPUTING↗

Easy_PERT: a Python tool for writing PERT cards and parsing PERT card results [Slides]

This presentation begins by providing an overview of the PERT card. The PERT card uses differential operator method to compute first- and second-order tally variations due to density, composition, and reaction cross-sections. It is possible to have multiple PERT cards in one MCNP input deck to study tally variations for several sets of nuclides, reactions, and energy ranges. Furthermore, the METHOD option tells MCNP to calculate either the perturbed tally (METHOD=-1, -2, -3) or the change in the unperturbed tally (METHOD=1, 2, 3). In summation, a powerful use-case for the MCNP code PERT card is that it facilitates calculating tally sensitivities to nuclear data. Writing PERT card entries and parsing output MCTAL files is tedious and error prone. however, Easy_PERT makes use of existing tools (Faust and MCNPTools) to handle writing PERT card entries and parsing the output MCTAL files. The PERT card is early in the development process and planned upcoming capabilities include calculating sensitivities and combining MCTAL files from separate runs into one JSON file.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Complex Parsing for In-Network Acceleration of High-Energy Physics Experiments

This paper describes a novel application and evaluation of programmable networking in High-Energy Physics (HEP): a complete parser for the custom packet format used by Fermilab’s DUNE experiment. Notably, this parser is implemented on a Tofino programmable network switch and evaluated on the FABRIC testbed by using network traffic generated by the ICEBERG DUNE prototype. The parsed network traffic consists of Jumbo Ethernet frames that contain digitizations of sensor readings from ICEBERG’s detector.This work is an early investigation into providing in-network processing support for HEP experiments. The paper describes DUNE’s custom packet format, the challenges encountered when implementing a parser for that format, and an exploration of the techniques that are needed to overcome those challenges. We identify performance bottlenecks and discuss directions for future research.

Sagstad, Bjoern [IIT, Chicago] (ORCID:000900033610↗

Toward a theory of distributed word expert natural language parsing

An approach to natural language meaning-based parsing in which the unit of linguistic knowledge is the word rather than the rewrite rule is described. In the word expert parser, knowledge about language is distributed across a population of procedural experts, each representing a word of the language, and each an expert at diagnosing that word's intended usage in context. The parser is structured around a coroutine control environment in which the generator-like word experts ask questions and exchange information in coming to collective agreement on sentence meaning. The word expert theory is advanced as a better cognitive model of human language expertise than the traditional rule-based approach. The technical discussion is organized around examples taken from the prototype LISP system which implements parts of the theory.

Rieger, C.↗

Information Extraction on an Earth Science Knowledge Graphs with Semantic Parsing

Knowledge graphs are an important tool, both for representing knowledge and for retrieving information. Fundamentally, they are semantic networks that represent entities and relationships in the form of nodes and edges. A large corpus of natural language text can bebroken down into discrete entities and relationships to form a useful knowledge graph. Existing research breaks down text into a subject, object, and verb relationship triple. Although this is a useful first step, it loses much of the original contextual information encoded within the text. Our process uses a novel 7-tuple approach, in which elements of sentences are programmatically parsed into seven categories: initiator, impacted, receiver, beneficiary, result, and context. In this presentation, we show a knowledge graph built using this 7-tupleprocessing of an Earth science corpus. We explain the techniques used to create the graph and analyze its information retrieval capability while assessing the accuracy and limitations of the results.

Carson Davis↗

Parsing-based Approaches for Verification and Recognition of Hierarchical Plans

Hierarchical Task Networks were proposed as a method todescribe plans by decomposition of tasks to sub-tasks untilprimitive tasks, actions, are obtained. Plan verification assumesa complete plan as input, and the objective is findinga task that decomposes to this plan. In plan recognition, aprefix of the plan is given and the objective is finding a taskthat decomposes to the (shortest) plan with the given prefix.This paper describes how to verify and recognize plans usinga common method known from formal grammars, by parsing.

Bartak, Roman↗

A Novel Parsing-based Approach for Verification of Hierarchical Plans

Hierarchical Task Networks were proposed as amethod to describe plans by decomposition of tasks to subtasksuntil primitive tasks, actions, are obtained. Valid plans –sequences of actions – must adhere both to causal dependenciesbetween the actions and to the structure given by the decompositionof the goal task. Plan verification aims at finding if a givenplan is valid, that is, if it is causally consistent and it can beobtained by decomposition of some task. The paper describes anovel parsing-based approach for hierarchical plan verificationthat is orders of magnitude faster than existing methods.

Bercher, Pascal↗

Auto Procedure Parsing: A Natural Language Processing Approach

Nuclear Power Plant (NPP) operating procedure is “a set of rules that describes how actions on the plant should be made if a certain system goal should be accomplished. U.S. NPPs use paper-based procedures (PBPs). PBPs are difficult to use. Common errors with PBPs are: following the wrong procedure, omit a step etc. Computer based procedures (CPBs) offer great improvement in ensuring plant safety. Some studies rely on experienced operators to understand the procedure content and then reorganize the procedure with digitally executable capabilities. Other studies utilize the procedure format to design rules to extract information from the operating procedures. Existing studies in procedure parsing are manual, laborious and extract limited information. This study aims to automatically extract critical information from operating procedures for generating computer interpretable representation of procedures. Such representation can further be used for automatic dynamic human reliability analysis and CPB design etc.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

TopTemp: Parsing Precipitate Structure from Temper Topology

Technological advances are in part enabled by the development of novel manufacturing processes that give rise to new materials or material property improvements. Development and evaluation of new manufacturing methodologies is labor-, time-, and resource-intensive expensive due to complex, poorly defined relationships between advanced manufacturing process parameters and the resulting microstructures. In this work, we present a topological representation of temper (heat-treatment) dependent material micro-structure, as captured by scanning electron microscopy, called TopTemp. We show that this topological representation is able to support temper classification of microstructures in a data limited setting, generalizes well to previously unseen samples, is robust to image perturbations, and captures domain interpretable features. The presented work outperforms conventional deep learning baselines and is a first step towards improving understanding of process parameters and resulting material properties.

Kassab, Lara↗

Tutorial: MATLAB Implementation of a Successive Convexification Algorithm for 3 DoF Rocket Landings

The primary objective of this work is to fill in gaps and explore an alternate way of solving the 3 DoF rocket-powered landing problem presented in the 2016 AIAA paper by Szmuk, Ackimese, and Berning using successive convexification (SCvx). In the original paper, CVX, an automatic parsing package, was used to transcribe the high-level trajectory optimization problem into a format that could be read by a conic solver. The parsing step, generally computationally intensive, is hidden from the user. The use of CVX is sufficient for the generation of trajectories off-line due to the lack of runtime and flight software implementation constraints. For on-line applications, it is necessary to parse the problem for flight software implementation. References on hand-parsing powered descent guidance (PDG) problems are sparse. In this Tech Memo, the process of transcribing the 3 DoF PDG problem into the format required by MATLAB’s built-in second-order cone solver, coneprog.m, is presented in detail. Due to the abridged 3 DoF dynamics and the relatively simple nonlinearities, this reference is the natural starting point for anyone interested in grasping the concepts behind SCvx pertaining to PDG and the parsing step. Simulation results shown in this report were independently created by solving the problem using coneprog.m. The intent of this memo is to serve as a supplemental material to the original paper by breaking down the concept behind successive convexification and shed light into the parsing process. Readers are encouraged to first familiarize themselves with the material laid out in the original reference.

Alex Hayes↗

EV-ELM (Electric Vehicle Policies with the Energy Language Model) [SWR-25-156]

Electric Vehicle Policies with the Energy Language Model (EV-ELM) leverages previous work using Large Language Models (LLMs) to find, download, and parse policy information related to energy infrastructure. In this application, we use LLMs to find policy documents related to the permitting and installation of electric vehicle charging infrastructure. This software contains the code to find, download, and parse these documents, while a related data record in the Open Energy Data Initiative (OEDI) will include the resulting output dataset that can be used for downstream analysis. The EV-ELM repository contains code for the EV-ELM project, which focuses on retrieving and processing EV permitting processes using large language models. The project is composed of two pipelines: (1) a web scraping pipeline for discovering and downloading EV permitting documents, and (2) a document parsing and extraction pipeline that processes the downloaded files to produce structured data. The web scraping pipeline is designed to extract relevant information from various websites, while the document parsing pipeline processes and analyzes the extracted documents to derive meaningful insights. Both pipelines depend on the NLR elm repository, which provides essential tools and functionalities for handling and processing the data. The web scraping pipeline is a modified version of the ordinance_gpt example within the elm repository. It has been adapted to fit the specific requirements of the EV-ELM project, ensuring that it effectively captures and processes the necessary information related to EV permitting.

Olson, Reid [National Laboratory of the Rockies (N↗

HDM/PASCAL Verification System User's Manual

The HDM/Pascal verification system is a tool for proving the correctness of programs written in PASCAL and specified in the Hierarchical Development Methodology (HDM). This document assumes an understanding of PASCAL, HDM, program verification, and the STP system. The steps toward verification which this tool provides are parsing programs and specifications, checking the static semantics, and generating verification conditions. Some support functions are provided such as maintaining a data base, status management, and editing. The system runs under the TOPS-20 and TENEX operating systems and is written in INTERLISP. However, no knowledge is assumed of these operating systems or of INTERLISP. The system requires three executable files, HDMVCG, PARSE, and STP. Optionally, the editor EMACS should be on the system in order for the editor to work. The file HDMVCG is invoked to run the system. The files PARSE and STP are used as lower forks to perform the functions of parsing and proving.

Hare, D.↗